flâneur — a map of the web's best reading

High-level visual representations in the human brain are aligned with large language models | Nature Machine Intelligence

nature.com · 250 words · saved by 1 readers

Doerig, Kietzmann and colleagues show that the brain’s response to visual scenes can be modelled using language-based AI representations. By linking brain activity to caption-based embeddings from large language models, the study reveals a way to quantify complex visual understanding.

Download PDF Subjects Cognitive neuroscience Neural encoding A preprint version of the article is available at arXiv. Abstract The human brain extracts complex information from visual inputs, including objects, their spatial and semantic interrelations, and their interactions with the environment. However, a quantitative approach for studying this information remains elusive. Here we test whether the contextual information encoded in large language models (LLMs) is beneficial for modelling the complex visual information extracted by the brain from natural scenes. We show that LLM embeddings of

Explore this link on the map →

saved by

related reading