Artificial neural networks (ANNs) can predict neural responses to natural images with remarkable accuracy, making them promising models of human vision. However, prediction scores alone reveal little about what computations these models have captured or how well they generalize beyond the images on which they are trained. Here, we introduce a principled zero-shot model evaluation framework in which ANN-based encoding models are frozen before being tested on not just held out images but entirely new datasets. Our strongest tests repurpose decades of cognitive neuroscience experiments as diagnostic benchmarks for identifying which computations current models capture and where they fail. To implement this framework we built encoding models of human category-selective regions (FFA, PPA, and EBA) using several pretrained ANNs and evaluated these fixed models across 7 independent fMRI datasets and 31 cognitive experiments from 10 published studies. This framework revealed three broad patterns. First, zero-shot evaluations exposed systematic limits to model predictivity that were largely hidden by standard within-dataset cross-validation. These limits were structured varying across brain regions and model classes and depending strongly on the similarity between the images used to build and test the models. Second, current models reproduced some, but not all, classical findings from cognitive neuroscience, providing a diagnosis of the computations that current models have yet to capture. Finally, models that performed well on prediction tests also tended to succeed at reproducing cognitive neuroscience findings, suggesting that prediction and explanation are closely linked. Unlike prediction scores however, cognitive tests help diagnose where and why models fail. Together, the zero-shot evaluation framework provides a scalable and unified approach for evaluating brain models across prediction and cognitive neuroscience tests, which can help us better understand what current models capture and what still remains to be explained.
Wang, R., Deb, M., Abate, A., Dipani, A., Dudipala, K. R., Chillarege, S., Ravikanti, K., Li, Y., Al-Tahan, H., Koushik, R., Mieczkowski, E., Fung, H., Kanwisher, N., Ratan Murty, N. A.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 0
- Comments 0
