NHK Developing Scene-adaptive Camera Sensors
Regular readers may have noted our report from the 2024 IBC show in Amsterdam that covered a new technology development from NHK Research. The organisation is working on a concept, so far just in a low resolution prototype, that is a camera sensor that can use what it calls ‘scene-adaptive’ imaging technology. The sensor, currently with just 1K resolution and with 272 ‘zones’, can capture each zone with different parameters – such as resolution, frame rate, light sensitivity and brightness.Â

This is quite a revolutionary idea that could significantly alter the way that cameras operate. At present, a videographer or photographer (let’s say a creator) often has to make a decision about the overall scene, for example on brightness. Current sensors have a certain dynamic range – that is to say the ratio between the brightest part of the image and the dark parts. Too much exposure and color and detail is lost in the bright parts of the scene, while with too little, darker details in the scene can be ‘crushed’ or masked by noise. Technologies such as log-based image capture and deeper bit depths can help, but in the end some compromise is often needed to get close to what they eye can see. The human eye has a great ability to adapt to different conditions. Â
Equally, creators sometimes want to get a ‘film look’ using, for example, a low frame rate. However, fast motion can then be difficult to show clearly (and Hollywood has to work hard to make fast motion look good in its 24fps film format). However, in the NHK approach, a creator could use a 24fps overall frame rate to give a ‘film look’, while showing a fast moving object (such as a ball or puck in a sports event) at a high frame rate within the image.

In a video (Japanese sound, but English subtitles), NHK showed a scene including a bright necklace, a dark model helicopter, a rotating object and an area of fine print (part of the image is in the photo above from IBC). An exposure optimized for any one part of the scene would lead to over- or under-exposure or the blurring of objects. Optimizing for each area would overcome these problems. An article from the researchers includes comparisons of image quality for the different parts of the image.
After Image Capture
Capturing an image is, of course, just one aspect of getting it to the viewer. In between, there is a lot of complexity, especially in video. Arguably, codecs do some of the same kinds of things as the variable sensor, encoding and sending different levels of data according to the content. However, they tend to use a single refresh rate and fixed resolution formats. On the other hand, as we have previously reported, NHK has been working on exploiting the feature in the VVC codec of allowing ‘enhancement layers’ that can be combined with a base image. We could imagine a scenario where a base fixed format stream is accompanied by an enhancement stream that adds the additional frames, resolution or dynamic range. (This could be particularly attractive if you combine traditional broadcast for the base layer with an enhancement layer over a different channel, i.e. by streaming).Â
Optimizing the amount of data to send resolution, frame rates or high bit depths only where really needed should allow the highest level of image quality for the bandwidth available over any particular distribution mechanism. That is, effectively, what a codec tries to do, especially a more recent one like VVC, where it can be combined with content dependent encoding, as Spin Digital and Intel have demonstrated.
Displays are a Challenge
When it comes to the display, variable resolution can be a challenge. Human vision has very high acuity in the central (foveal) part of the eye, but less elsewhere. This property has been exploited, for example in head-mounted displays from companies such as Varjo of Finland that developed high end systems to merge a high resolution center display with a lower resolution peripheral display. It’s difficult to do, even for a single viewer and when you know where their eyes are and where they are looking, but not practical for multiple viewers. This is where 8K can be a real advantage, providing a high level of resolution on all of the image allowing the viewer to focus on the detail.Â
Display makers are already somewhat familiar with this topic as it is similar to the way that zones are used in miniLED TVs where the display is divided into different zones and LCD backlights are adjusted appropriately for each zone. Where a deep black is needed, the zone is dimmed, while for highlights it is boosted. However, actually displaying different parts of the image at different frame rates would need some fairly radical changes to the way that flat panel displays are driven. On the other hand, it’s not tricky to slow down a frame rate, using repeated frames, so just as the resolution of an area can be high if the whole display has high resolution, the frame rate can be high if all the display is refreshed at high frequency.
One of the challenges for display makers with this kind of variability is that while camera sensors are often built using CMOS semiconductor technology, allowing for sophisticated processing within the sensor device, the transistors used in displays are typically of low quality and speed in comparison, so it is hard to do a lot of processing at the individual pixel level (although over the years, there have been several attempts to develop displays with processing attached to each pixel).
Onwards to 32K!
At IBC, NHK told us that it regards 8K as ‘done’ from an R&D point of view. However that is not the end of the road for resolution. The organization told us that the ITU is looking to 32K video with 120 fps as a target for 360 degree immersive video, so research is continuing as we’re a long way from that at the moment!
