Premium accounts now available! Sign up and create a premium account. Read more Close

Advertisement

Image

Behaviorally prioritized entity-relation structure captures human visual cortical representations of natural scenes

Preprint Created on 18 Sep 2026 bioRxiv

Understanding natural scenes requires identifying visible entities and representing how those entities are related. Recent studies have shown that artificial neural networks (ANNs), large language models (LLMs), and vision language models (VLMs) can predict visual cortical responses to natural images. However, the neural organization of relational scene meaning remains poorly understood, in part because these models typically encode scene content in global feature spaces that are difficult to decompose into separable entity and relation components. Here, we combined scene-graph annotations, behavioral measurements, and large-scale neural datasets to characterize structured relational representations during natural vision. We used RotatE, a knowledge-graph embedding model, to represent head-relation-tail triplets annotated for images from the 7T Natural Scenes Dataset. Triplet embeddings reliably captured cortical representational structure across the visual hierarchy. Behavioral judgments further revealed systematic differences in triplet accessibility associated with visual, relational, and graph properties. Prioritizing more behaviorally accessible triplets improved neural correspondence and explained unique variance beyond object co-occurrence, ANN image features, and LLM caption embeddings. Decomposing triplet representations into entity and relation components revealed partially dissociable cortical contributions, with lateral parietal cortex showing sensitivity to both. Triplet-based semantic information also remained spatially grounded: visual-field-specific triplet models preferentially predicted voxels with matching retinotopic preferences. Finally, cross-species comparison indicated that triplet-based semantic features were relatively more aligned with human high-level visual cortex than with macaque inferotemporal cortex. Together, these findings provide new insights into the representation of semantic relational information in the human visual cortex during natural scene perception.

Wu, Y., Jiang, W., Li, S.

Advertisement

Stats

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 2
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement