NTH

On Locality and Length Generalization in Visual Reasoning

AuthorsPulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya, Roland Memisevic

July 29, 2026 2 min read
Watch on YouTube
The one-line take

This study finds that vision models relying on local, sequential glimpses can generalize better to longer and more complex visual reasoning tasks than models using global shortcuts.

Key results

32.63
Finding Roots global baseline OOD-numroots

Accuracy (%) of global-view Qwen2.5-VL-3B-Instruct on the out-of-distribution number-of-roots split.

67.12
Finding Roots foveated OOD-numroots

Accuracy (%) of FoveAgent-Qwen with a 1200x800 global glimpse plus high-resolution local glimpses.

29.0%
Foveated matched-compute gain

Accuracy improvement attributed to FoveAgent-Qwen over a compute-matched global-view baseline.

3.8%
Global compute scaling gain

Accuracy improvement from scaling global visual compute by 10x.

What the paper found

Researchers at Qualcomm AI Research study why vision models fail to generalize when visual problems become longer or more complex than their training examples. They introduce four tasks—Visual Parity, State Machine, Recall, and Finding Roots—where relevant evidence is distributed across images, and compare global-view systems including Qwen2.5-VL-3B-Instruct, GPT-5.4 from OpenAI, and Claude Sonnet 4.6 from Anthropic. Their central model, FoveAgent-LSTM, combines a high-resolution local glimpse, a low-resolution peripheral glimpse, and recurrent LSTM state updates; unlike global transformers, it learns a step-by-step procedure that extrapolates to more switches, larger canvases, and longer visual dependencies. Controlled experiments show that recurrence alone is insufficient: recurrent models with full-image access also learn global shortcuts, while strict locality prevents those shortcuts. On the real-world-inspired Finding Roots task from Math-Search, a global Qwen baseline reaches 32.63 percent on the out-of-distribution number-of-roots split, compared with 67.12 percent for FoveAgent-Qwen using local crops. Under matched visual compute, foveated processing adds 29.0 percent accuracy, whereas scaling global compute by 10x adds only 3.8 percent. The contrast with Recall is important: global models handle cluttered retrieval better, indicating that local recurrent perception is specifically valuable for compositional state tracking, not every form of visual search. The paper’s conclusion is that robust visual length generalization requires both local information gathering and strictly recurrent computation.

Original abstract

A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular computer vision models in use today, which input images globally and in a single shot. A natural question therefore is whether local, sequential vision models may provide any fundamental computational benefits in addition to being biologically more plausible than global models. In this work, we investigate this question from the perspective of visual state tracking and length generalization. Inspired by recent studies of length generalization in language models, we study the behavior of vision models trained on simple vision tasks that require the aggregation of local information across an image. Our experiments reveal that, similar to language models, vision models can learn to exploit global shortcuts and thereby fail to generalize over task length or complexity. We also show that recurrent vision policies based on strictly local perception can mitigate these failures, thereby allowing models to generalize on these tasks. Our results show that local attention may be an essential overlooked requirement for robust compositional generalization.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis