Thought without systematicity? Evaluating reasoning models on rule induction tasks
AuthorsSimon Schug, Brenden M. Lake
AffiliationsPrinceton University
Resources
Reasoning models can solve individual rule tasks yet often fail when the same underlying logic is rearranged, suggesting their thinking is less systematic than it appears.
Key results
Number of structurally equivalent variations sampled per task.
Program-induction variation sets where Claude Opus 4.7 solved at least one variant.
Program-induction variation sets where Claude Opus 4.7 solved every variant.
Difference between Claude Opus 4.7 solving any versus all program-induction variants.
What the paper found
The paper tests whether reasoning models exhibit systematicity: the ability to preserve task-solving competence when symbols, feature bindings, example order, or constituent arrangements change without altering the underlying rule. It constructs structurally equivalent variants of four synthetic rule-induction families—grammar-based instruction-learning, symbolic Raven’s progressive matrices, program induction over integer sequences, and Boolean category learning—and evaluates OpenAI’s GPT-5.5, Anthropic’s Claude Opus 4.7, and Google’s Gemini 3.1 models. Each task is expanded into K=5 isomorphic variations, with each variation judged by majority vote over N=5 sampled attempts. The central result is a systematicity gap: models often solve at least one variant but fail on equivalent variants, especially in Raven’s matrices and integer-sequence program induction. For Claude Opus 4.7, 91.67% of program-induction variation sets had at least one solved variant, but only 50.00% had all variants solved, a 41.67% gap. Boolean category learning was the exception, with almost all models solving nearly every variation. Increasing reasoning effort from low to high did not consistently improve systematicity, and Gemini’s zero-temperature evaluation showed that the gap persists even when sampling noise is largely removed. The findings challenge the assumption that success on a single reasoning benchmark reliably indicates a transferable underlying capability.
Original abstract
A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms such as recombination and substitution. We find that despite being able to correctly solve a task, models often fail on structurally equivalent variants of the same task. These findings suggest that many model behaviors lack systematicity, rendering it difficult to robustly establish the cognitive abilities of reasoning models beyond the particular contexts they were evaluated in.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.