Researchers in artificial intelligence increasingly intervene on language models to study how their functions are organized, silencing weights and attention components in ways reminiscent of the brain lesions long used to map language in post-stroke aphasia. Under this program of mechanistic interpretability, a common move is to ablate an attention head and read the resulting drop in a behavior as evidence that the head implements it. How much that move can establish about where a behavior is computed remains unclear, because a head whose removal disrupts a behavior is necessary for it but need not be the site where it is computed. We examined this question for the picture naming task, taking the loss of naming (anomia) that defines aphasia as the behavior of interest. Across six vision-language models, spanning three language backbones and a range of parameter scales, we ablated each attention head in turn during single-word picture naming and measured accuracy before and after. The degree of localization varied widely across models. In LLaVA-1.6-Vicuna-13B, a single early head (layer 0, head 20) was necessary: removing it alone reduced naming accuracy from 0.99 to 0.006. The same head was not sufficient, because retaining it while ablating the other 1{,}599 heads also produced 0% accuracy. Two Mistral-backbone models (LLaVA-Mistral-7B and Idefics2-8B) had no critical head. An early-layer dependence was present in every model but varied in strength, whereas dependence on any single head ranged from dominant to absent. Within Qwen2.5-VL, a dominant head was present in the 7B model but not the 3B model, indicating that this concentration emerged with scale rather than being fixed across a model family, and it was not explained by attention type. These results show that ablating the head whose removal disrupts naming does not establish that the head computes the behavior, that the result generalizes across models, or that the behavior localizes to a head at all.
Nemati, S., Newman-Norlund, R. D., Ahmadi, S., Guan, X., Warren, K., Yang, Y., Nelakuditi, S., Rorden, C., Bonilha, L., Fridriksson, J.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 0
- Comments 0
