On Locality and Length Generalization in Visual Reasoning
AuthorsPulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya, Roland Memisevic
Resources
This study finds that vision models relying on local, sequential glimpses can generalize better to longer and more complex visual reasoning tasks than models using global shortcuts.
Key results
Accuracy (%) of global-view Qwen2.5-VL-3B-Instruct on the out-of-distribution number-of-roots split.
Accuracy (%) of FoveAgent-Qwen with a 1200x800 global glimpse plus high-resolution local glimpses.
Accuracy improvement attributed to FoveAgent-Qwen over a compute-matched global-view baseline.
Accuracy improvement from scaling global visual compute by 10x.
What the paper found
Researchers at Qualcomm AI Research study why vision models fail to generalize when visual problems become longer or more complex than their training examples. They introduce four tasks—Visual Parity, State Machine, Recall, and Finding Roots—where relevant evidence is distributed across images, and compare global-view systems including Qwen2.5-VL-3B-Instruct, GPT-5.4 from OpenAI, and Claude Sonnet 4.6 from Anthropic. Their central model, FoveAgent-LSTM, combines a high-resolution local glimpse, a low-resolution peripheral glimpse, and recurrent LSTM state updates; unlike global transformers, it learns a step-by-step procedure that extrapolates to more switches, larger canvases, and longer visual dependencies. Controlled experiments show that recurrence alone is insufficient: recurrent models with full-image access also learn global shortcuts, while strict locality prevents those shortcuts. On the real-world-inspired Finding Roots task from Math-Search, a global Qwen baseline reaches 32.63 percent on the out-of-distribution number-of-roots split, compared with 67.12 percent for FoveAgent-Qwen using local crops. Under matched visual compute, foveated processing adds 29.0 percent accuracy, whereas scaling global compute by 10x adds only 3.8 percent. The contrast with Recall is important: global models handle cluttered retrieval better, indicating that local recurrent perception is specifically valuable for compositional state tracking, not every form of visual search. The paper’s conclusion is that robust visual length generalization requires both local information gathering and strictly recurrent computation.
Original abstract
A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular computer vision models in use today, which input images globally and in a single shot. A natural question therefore is whether local, sequential vision models may provide any fundamental computational benefits in addition to being biologically more plausible than global models. In this work, we investigate this question from the perspective of visual state tracking and length generalization. Inspired by recent studies of length generalization in language models, we study the behavior of vision models trained on simple vision tasks that require the aggregation of local information across an image. Our experiments reveal that, similar to language models, vision models can learn to exploit global shortcuts and thereby fail to generalize over task length or complexity. We also show that recurrent vision policies based on strictly local perception can mitigate these failures, thereby allowing models to generalize on these tasks. Our results show that local attention may be an essential overlooked requirement for robust compositional generalization.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.