NTH

CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning

AuthorsLinas Nasvytis, Simon Jerome Han, Ben Prystawski, Satchel Grant, Noah D. Goodman, Judith E. Fan

July 15, 2026 2 min read
Watch on YouTube
The one-line take

CORE helps language models get better at reasoning faster by turning comparisons between right and wrong attempts into compact, human-readable strategy notes.

Key results

350
Rapid improvement checkpoint

At 350 training rollouts, CORE exceeded the best baseline performance at any training point.

59.9%
Average accuracy improvement

Average held-out accuracy increased from 0.445 to 0.712.

9
Best task-regime results

CORE achieved the highest mean accuracy in 9 of 12 task-and-training-data conditions.

0.92k
Evaluation context

Average additional tokens per evaluation item, compared with 33.6k for Episodic RAG and 32.7k for MemRL.

0.268
Contrastive reflection ablation

Final mean held-out accuracy improvement for full CORE across the four tasks.

What the paper found

Researchers at Stanford University introduce CORE, or Contrastive Reflection, a non-parametric method for improving a frozen language model from verifier rewards without updating its weights. Using OpenAI’s gpt-oss-120b, CORE maintains separate memories for successful reasoning traces and concise natural-language insights. After a failed attempt, it contrasts that trace with a semantically similar successful trace, proposes strategies or constraints, tests each insight against the verifier, and retrieves useful insights later using both semantic similarity and empirical utility. Across Tower of Hanoi, MathGAP, ZebraLogic, and Matchstick Arithmetic, CORE reached stronger performance than GRPO, GEPA, Episodic RAG, and MemRL with fewer rollouts: at 350 rollouts, it exceeded the best baseline result at any training checkpoint, and average held-out accuracy improved by 59.9%, from 0.445 to 0.712. With only 5, 10, or 100 training problems, CORE achieved the best result in 9 of 12 task-and-data conditions. It also added just 0.92k evaluation-time tokens per item, versus 33.6k for Episodic RAG and 32.7k for MemRL. Ablations show that contrasting success with failure produced a final mean accuracy improvement of 0.268, compared with 0.227 when utility-aware retrieval was removed. The resulting memories are compact, interpretable reasoning abstractions, although CORE requires verifiable rewards and incurs extra inference cost for reflection and admission testing.

Original abstract

Language models can use verifiable rewards to improve at a wide variety of reasoning tasks. However, both parametric (e.g. RLVR) and non-parametric (e.g. prompt optimization) approaches to doing so typically require hundreds of training samples and thousands of model rollouts, making them expensive in the best case and intractable in the worst. To address this challenge, we introduce Contrastive Reflection (CORE), a non-parametric learning algorithm that compares past reasoning traces to generate insights: short natural-language descriptions of reasoning strategies and constraints that capture differences between successful and unsuccessful problem attempts. Across four reasoning tasks, we demonstrate that CORE enables more rapid improvement than both parametric (GRPO) and non-parametric (GEPA, episodic RAG, and MemRL) methods, while using fewer rollouts. Under fixed rollout budgets with as few as five training samples, we then show that CORE also achieves comparable or greater performance gains than each baseline. Finally, we highlight how CORE is also substantially more context-efficient than non-parametric baselines, requiring fewer prompt tokens while storing learned knowledge as compact, interpretable natural-language insights. Our results therefore suggest that distilling contrasts between successful and unsuccessful reasoning traces into abstract and useful insights can provide a more efficient and interpretable route to model self-improvement than weight updates, prompt optimization, or direct reuse of stored reasoning traces.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis