NTH

Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate

AuthorsJohn Seon Keun Yi, Aaron Mueller, Dokyun Lee

June 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how to compress expensive multi-agent debate into one language model, making reasoning cheaper and exposing interpretable internal agent-like directions that can also be steered for safety control.

Key results

944
Debate traces

The IMAD pipeline is trained on 944 structured three-agent, two-round debate traces generated with GPT-3.5-turbo.

85.20%
GSM8K accuracy

On LLaMA-3.1 8B, IMAD (SFT+RL) reaches 85.20% on GSM8K.

15.41%
ROUGE-L faithfulness improvement

Agent-specific steering on the internalized model improves ROUGE-L faithfulness AUC by 15.41% on average over the base model.

What the paper found

Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate proposes IMAD, a two-stage post-training pipeline that distills explicit multi-agent debate into a single LLM by first supervised fine-tuning on full debate transcripts and then applying GRPO reinforcement learning with a decaying format reward and an annealed length-clipping reward. Trained on 944 three-agent, two-round arithmetic debates generated with GPT-3.5-turbo and evaluated on LLaMA-3.1 8B, Qwen 2.5 7B, and Mistral Nemo 12B, IMAD matches or exceeds explicit debate on GSM8K, MMLU-Pro, and Big-Bench Hard while reducing token usage to 6.3%–21.1% of Debate, a 5–16× inference efficiency gain; on LLaMA-3.1 8B it reaches 85.20% on GSM8K versus 83.03% for Debate. The paper’s mechanistic contribution is that internalization preserves agent identity as separable activation subspaces: contrastive activation addition and difference-in-means steering produce agent-specific directions whose ROUGE-L faithfulness AUC improves by 15.41% on average over the base model, with the strongest gain for the Program-of-Thought agent. This structure also enables safer control: when the authors internalize deliberately malicious “evil” and hallucinating agents, negative steering suppresses harmful behavior more cleanly than on the base model, reaching near-zero evil scores while preserving GSM8K accuracy and lowering perplexity, suggesting that multi-agent debate can be compressed into latent, controllable reasoning circuits rather than verbose transcripts.

Original abstract

Multi-agent debate has been shown to improve reasoning in large language models (LLMs). However, it is compute-intensive, requiring generation of long transcripts before answering questions. To address this inefficiency, we develop a framework that distills multi-agent debate into a single LLM through a two-stage fine-tuning pipeline combining debate structure learning with internalization via dynamic reward scheduling and length clipping. Across multiple models and benchmarks, our internalized models match or exceed explicit multi-agent debate performance using up to 93% fewer tokens. We then investigate the mechanistic basis of this capability through activation steering, finding that internalization creates agent-specific subspaces: interpretable directions in activation space corresponding to different agent perspectives. We further demonstrate a practical application: by instilling malicious agents into the LLM through internalized debate, then applying negative steering to suppress them, we show that distillation makes harmful behaviors easier to localize and control with smaller reductions in general performance compared to steering base models. Our findings offer a new perspective for understanding multi-agent capabilities in distilled models and provide practical guidelines for controlling internalized reasoning behaviors. Code available at https://github.com/johnsk95/latent_agents

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis