NTH

Thought without systematicity? Evaluating reasoning models on rule induction tasks

AuthorsSimon Schug, Brenden M. Lake

AffiliationsPrinceton University

September 23, 2026 2 min read
Watch on YouTube
The one-line take

Reasoning models can solve individual rule tasks yet often fail when the same underlying logic is rearranged, suggesting their thinking is less systematic than it appears.

Key results

5
Isomorphic task variations

Number of structurally equivalent variations sampled per task.

91.67%
Claude program induction solve-any rate

Program-induction variation sets where Claude Opus 4.7 solved at least one variant.

50.00%
Claude program induction solve-all rate

Program-induction variation sets where Claude Opus 4.7 solved every variant.

41.67%
Program induction systematicity gap

Difference between Claude Opus 4.7 solving any versus all program-induction variants.

What the paper found

The paper tests whether reasoning models exhibit systematicity: the ability to preserve task-solving competence when symbols, feature bindings, example order, or constituent arrangements change without altering the underlying rule. It constructs structurally equivalent variants of four synthetic rule-induction families—grammar-based instruction-learning, symbolic Raven’s progressive matrices, program induction over integer sequences, and Boolean category learning—and evaluates OpenAI’s GPT-5.5, Anthropic’s Claude Opus 4.7, and Google’s Gemini 3.1 models. Each task is expanded into K=5 isomorphic variations, with each variation judged by majority vote over N=5 sampled attempts. The central result is a systematicity gap: models often solve at least one variant but fail on equivalent variants, especially in Raven’s matrices and integer-sequence program induction. For Claude Opus 4.7, 91.67% of program-induction variation sets had at least one solved variant, but only 50.00% had all variants solved, a 41.67% gap. Boolean category learning was the exception, with almost all models solving nearly every variation. Increasing reasoning effort from low to high did not consistently improve systematicity, and Gemini’s zero-temperature evaluation showed that the gap persists even when sampling noise is largely removed. The findings challenge the assumption that success on a single reasoning benchmark reliably indicates a transferable underlying capability.

Original abstract

A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms such as recombination and substitution. We find that despite being able to correctly solve a task, models often fail on structurally equivalent variants of the same task. These findings suggest that many model behaviors lack systematicity, rendering it difficult to robustly establish the cognitive abilities of reasoning models beyond the particular contexts they were evaluated in.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis