Back to Home
Technology

Study reveals what the brain understands that artificial intelligence models cannot

A study based on behavioural experiments and magnetic resonance brain imaging has shown that humans perceive places through the opportunities for movement and action they offer, rather than through their visual features alone. A comparison found that artificial intelligence networks trained to classify images do not replicate this ability efficiently because they lack embodied sensorimotor experience.

•7 min

Listen to this article

An automatically generated audio version.

0:00
0:00
An illustration of a human brain glowing blue, surrounded by waveforms and lines suggesting brain activity.

A study published in the journal PNAS has found that the human brain does not merely recognise the components of a place when viewing it, but automatically builds a representation of the activities and movements possible within it. Image-, scene- and video-trained artificial intelligence models have been unable to replicate this ability with the same efficiency.

Action possibilities explain how humans perceive places

The study examined the concept of affordances, proposed by American psychologist James Gibson in the 20th century. It holds that people perceive places through what they can do in them, rather than solely through their shapes or visible features. When looking at a beach, for example, the brain sees more than sand and water; it also treats it as a place that affords swimming or building sandcastles.

Mai Mabrouk, a professor of biomedical informatics at the Faculty of Information Technology and Computer Science at Egypt’s Nile University, who was not involved in the study, explains that the difference between seeing a beach as a place for swimming and seeing it as merely sand and water is that the brain does not simply describe the scene, but derives from it the possibility of moving within it.

She added that this is a neural process distinct from merely recognising the elements or materials present in an image. The researchers used behavioural experiments and functional magnetic resonance imaging to record brain activity.

231
صورة لأماكن متنوعة

They showed a group of volunteers 231 images of diverse places, including indoor and outdoor environments as well as natural and built settings. Participants were then asked to identify the actions they could perform in each environment shown. The movement options they assessed included walking, cycling, swimming, driving, climbing and sailing.

The task was to infer the possible action from each image, allowing the researchers to study the relationship between the scene’s visual features and the movement opportunities people perceive in it. While the images were being viewed, the team used functional magnetic resonance imaging to monitor the volunteers’ brain activity and identify the regions that responded to scenes affording particular types of movement.

Brain activity is linked to movement opportunities in the environment

For the comparison stage, the researchers used neural networks trained to classify images to test how well they could represent a scene in the way the human brain does. The results showed strong activity in specific areas of the participants’ brains when they viewed certain environments. This activity was linked not only to the content of the images, but also to the movement possibilities they offered.

When seeing the sea, for example, the brain responded to the possibility of swimming offered by the place, rather than merely to the presence of water in the image. The finding indicates that the brain does not stop at recognising visual elements such as rocks, floors or other components of a scene, but automatically creates representations of what can be done in each environment.

These possibilities may include actions such as walking or climbing, as well as other movement-related outcomes such as falling. The study identified two specialised visual regions associated with this motor representation: the parahippocampal place area and the occipital visual area. Representations of affordances are built even without directing a person’s attention towards them or assigning a specific perceptual task, meaning that the brain maps movement simply by looking at the environment.

Mabrouk said the brain continuously creates motor maps, and that these maps are essential for immediate and effective interaction with the surroundings. She believes this mechanism supports action-oriented perception, since a person’s understanding of what they see cannot be separated from their readiness to move or make an appropriate decision within a place.

The experiments also showed that the neural representation of affordances does not depend on the type of task assigned to an individual. Whether participants were asked to describe the things they saw or simply to stare at a point, the brain regions remained active in a way indicating that the environment was perceived as an opportunity for a particular movement.

This means that representing movement opportunities is a fundamental component of human perception of the world, rather than a response that appears only when consciously thinking about carrying out an action. This characteristic may explain the ease with which people make movement decisions in unfamiliar environments, without needing to consciously analyse every object, surface or material in the scene.

Mabrouk added that the study showed that affordances are not derived solely from scene components such as surfaces and materials. Instead, the brain represents them within an independent representational space that is processed directly in the visual cortex. This representation separates knowledge of the objects present from an understanding of the motor functions afforded by the place.

The motor map identified by the study does not correspond to traditional scene classifications, such as open or enclosed environments. Nor is it based solely on the type of material, whether wood, stone or water, or on visible objects such as trees, buildings and streets. Instead, it takes the form of a three-dimensional representational map linking specific elements to the possible actions associated with them.

Examples of these links include associating the sea with the possibility of swimming, a boat with rowing and rocks with climbing. In this sense, the brain processes the functional relationship between an element and an action, rather than merely assigning a label to the element or a general classification to the place in which it appears.

Image models cannot replicate motor perception

When the researchers compared brain representations with the responses of artificial intelligence models trained to classify images, scenes or even video clips, they found a gap in the models’ ability to replicate this type of perception. Although deep neural networks are effective at recognising elements, they were weak at representing the motor affordances associated with them. Mabrouk said the study links this weakness to the models’ lack of embodied experience.

Artificial intelligence has no body that moves and interacts with the environment, so it does not receive the sensorimotor inputs that feed human action representations. Current models have also been trained primarily to label objects, rather than to use them or interact with them in practice.

Under this interpretation, it is not enough for a machine to recognise what it sees; it also needs experience linked to action within the environment. Mabrouk said narrowing the gap would require incorporating the body into sensorimotor learning, so that artificial intelligence learns by performing actions in the surrounding world, as mobile robots do.

She added that the crucial step was embodied learning in natural environments, which could enable models to develop functional representations closer to those produced by the human brain. Under this approach, learning would be linked to movement and sensory experience rather than relying solely on images and visual data.

Applications of embodied perception in driving and virtual reality

The findings open the way to practical applications, including improving the algorithms used by self-driving cars so they can better predict pedestrian behaviour and identify places that can be crossed. Virtual and augmented reality environments could also benefit from incorporating dynamic representations of movement possibilities, enhancing users’ interaction with digitally designed spaces.

The study also supports efforts to develop embodied artificial intelligence that learns from sensorimotor experience rather than from viewing images alone.

The neural map identified by the researchers shows that human perception is linked to action and decision-making, and that movement and the body are essential to the way people understand their surrounding environment.