This page is based on the event record supplied by the Biolà team. Paper links point to the publisher or DOI record.

About this conversation
Vision-language models and human lesion data provide converging evidence that language-related processing dynamically shapes visual representation.
We thank Haoyang Chen for sharing the research, the thinking behind the discovery, and the challenges and decisions that shaped the work with the Biolà community.
PAPERS & REFERENCES
Nature Human BehaviourEVENT MATERIALS
On January 11, 2026, from 9:00 to 10:00 AM Beijing Time, Biolà hosted the 18th issue of the Bioneer First-Author Forum, featuring Haoyang Chen, a PhD student at the School of Psychological and Cognitive Sciences, Peking University. The session was delivered in Chinese. As the first author of the featured study, Chen presented “Combined Evidence from Artificial Neural Networks and Human Brain-Lesion Models Reveals That Language Modulates Vision in Human Perception,” published in Nature Human Behaviour in 2025.
Comparing the internal representations of deep neural networks with human brain activity has become an important approach for investigating the computational principles of perception. Chen and colleagues examined why vision–language models such as CLIP show stronger correspondence with activity in the human ventral occipitotemporal cortex (VOTC) than conventional visual models. Across four datasets, CLIP consistently outperformed both the label-supervised ResNet and unsupervised MoCo models in predicting VOTC activity, with this advantage showing a left-hemisphere bias consistent with the lateralization of the human language system.
To move beyond correlational model–brain comparisons, the study further analyzed data from 33 patients with stroke-related brain lesions. Reduced white-matter integrity between the VOTC and the language-related left angular gyrus was associated with poorer CLIP–brain correspondence and, conversely, stronger correspondence with MoCo. These findings provide converging evidence that language-related processing dynamically modulates human visual representations. More broadly, the work demonstrates how brain-lesion models can offer causal constraints for evaluating artificial neural networks and developing computational models that more closely reflect human cognition.
We sincerely thank Haoyang Chen for sharing the findings and scientific reasoning behind this work with the Biolà community, and for discussing how artificial neural networks and human lesion data can be combined to investigate interactions between language and visual perception.
Citation: Chen, H., Liu, B., Wang, S., et al. (2025). Combined evidence from artificial neural networks and human brain-lesion models reveals that language modulates vision in human perception. Nature Human Behaviour. https://doi.org/10.1038/s41562-025-02357-5
