NTH
Research collection

AI Reasoning research

Explore methods and evaluations for mathematical, logical, and multi-step reasoning. Compare gains against inference cost and benchmark limitations.

39 papers · Latest edition October 5, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All AI Reasoning papers

Newest editions first.

02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis
06Reasoning

Towards a Deterministic Math Solver for Clinical Language Models

Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange, Rafi Al Attrach, Sahil Kapadia, Zakaria Laouabdia Sellami, Angelo Antonio Talio, Leo Anthony Celi

Instead of trusting language models to do clinical math, the system has them write code for a restricted executor—and shows that this helps larger models but is no universal cure for medical calculation errors.

Read analysis
08Reasoning

Fractal basins trap latent reasoning

Jeffrey Lai, Anthony Bao, John Quinn, William Gilpin

The study argues that when AI reasoning gets harder, models become temporarily chaotic and get trapped near almost-correct solutions.

Read analysis
09Reasoning

Thinking with Looped Flows

Ayhan Suleymanzade, Chanhyuk Lee, Floor Eijkelboom, Nicholas M. Boffi, İsmail İlkan Ceylan, Jinwoo Kim

Looped Flows lets models spend more computation at inference time by repeatedly refining noisy predictions, achieving strong results on several challenging reasoning benchmarks.

Read analysis
16Reasoning

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke

A powerful model can effectively teach a weaker one new problem-solving habits at test time by building a smart software harness around it.

Read analysis
20Reasoning

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary

This paper explains how to fairly compare reasoning LLMs that spend more computation at inference time and provides practical standards for evaluating and reproducing them.

Read analysis
26Reasoning

On Locality and Length Generalization in Visual Reasoning

Pulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya, Roland Memisevic

This study finds that vision models relying on local, sequential glimpses can generalize better to longer and more complex visual reasoning tasks than models using global shortcuts.

Read analysis
27Reasoning

Test-Time Scaling via Error Localization

Rajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta, Deepanway Ghosal, Aravindan Raghuveer

TTEL makes language-model reasoning more efficient by finding where a solution went wrong and branching from the last correct step instead of starting over.

Read analysis
32Reasoning

CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning

Linas Nasvytis, Simon Jerome Han, Ben Prystawski, Satchel Grant, Noah D. Goodman, Judith E. Fan

CORE helps language models get better at reasoning faster by turning comparisons between right and wrong attempts into compact, human-readable strategy notes.

Read analysis
33Reasoning

Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation

Joshua Ong Jun Leang, Zheng Zhao, Mihaela Cătălina Stoian, Qiyuan Xu, Haonan Li, Wenda Li, Shay B. Cohen, Eleonora Giunchiglia

This paper makes Lean theorem proving cheaper and stronger by combining curriculum training, proof-trace filtering, and a novel data augmentation scheme, while also testing a diffusion-style prover.

Read analysis
34Reasoning

A Verifiable Search Is Not a Learnable Chain-of-Thought

Harsh Patel

This paper argues that some reasoning problems can be solved by search but still cannot be learned as a neat left-to-right chain of thought—what models can distill is often verification and memorization, not the search itself.

Read analysis
36Reasoning

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

Jiacheng Chen, Xinyu Zhang, Shunkai Zhang, Yanmohan Wang, Lin Li, Tiancheng Qin, Qin Wang, Zhengmao Zhu, Tianle Li, Jingyang Li, Zehan Li, Binyang Jiang, Jin Zhu, Han Ding, Fei Yu, Chenyu Du, Zijian Song, Jiayuan Song, Zhi Zhang, Yunan Huang, Weiyu Cheng, Pengyu Zhao, Yu Cheng

MaxProof uses a population of proof candidates plus generator-verifier-ranker loops to push an AI system to gold-medal-level performance on hard math olympiad proofs.

Read analysis
37Reasoning

Self-Improving Language Models with Bidirectional Evolutionary Search

Guowei Xu, Zhenting Qi, Huangyuan Su, Weirui Ye, Himabindu Lakkaraju, Sham M. Kakade, Yilun Du

This paper introduces a new way for language models to improve themselves by both evolving candidate answers forward and breaking problems into checkable subgoals backward, boosting performance on hard reasoning tasks.

Read analysis