Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
AuthorsJohn Seon Keun Yi, Aaron Mueller, Dokyun Lee
This paper shows how to compress expensive multi-agent debate into one language model, making reasoning cheaper and exposing interpretable internal agent-like directions that can also be steered for safety control.
Key results
The IMAD pipeline is trained on 944 structured three-agent, two-round debate traces generated with GPT-3.5-turbo.
On LLaMA-3.1 8B, IMAD (SFT+RL) reaches 85.20% on GSM8K.
Agent-specific steering on the internalized model improves ROUGE-L faithfulness AUC by 15.41% on average over the base model.
What the paper found
Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate proposes IMAD, a two-stage post-training pipeline that distills explicit multi-agent debate into a single LLM by first supervised fine-tuning on full debate transcripts and then applying GRPO reinforcement learning with a decaying format reward and an annealed length-clipping reward. Trained on 944 three-agent, two-round arithmetic debates generated with GPT-3.5-turbo and evaluated on LLaMA-3.1 8B, Qwen 2.5 7B, and Mistral Nemo 12B, IMAD matches or exceeds explicit debate on GSM8K, MMLU-Pro, and Big-Bench Hard while reducing token usage to 6.3%–21.1% of Debate, a 5–16× inference efficiency gain; on LLaMA-3.1 8B it reaches 85.20% on GSM8K versus 83.03% for Debate. The paper’s mechanistic contribution is that internalization preserves agent identity as separable activation subspaces: contrastive activation addition and difference-in-means steering produce agent-specific directions whose ROUGE-L faithfulness AUC improves by 15.41% on average over the base model, with the strongest gain for the Program-of-Thought agent. This structure also enables safer control: when the authors internalize deliberately malicious “evil” and hallucinating agents, negative steering suppresses harmful behavior more cleanly than on the base model, reaching near-zero evil scores while preserving GSM8K accuracy and lowering perplexity, suggesting that multi-agent debate can be compressed into latent, controllable reasoning circuits rather than verbose transcripts.
Original abstract
Multi-agent debate has been shown to improve reasoning in large language models (LLMs). However, it is compute-intensive, requiring generation of long transcripts before answering questions. To address this inefficiency, we develop a framework that distills multi-agent debate into a single LLM through a two-stage fine-tuning pipeline combining debate structure learning with internalization via dynamic reward scheduling and length clipping. Across multiple models and benchmarks, our internalized models match or exceed explicit multi-agent debate performance using up to 93% fewer tokens. We then investigate the mechanistic basis of this capability through activation steering, finding that internalization creates agent-specific subspaces: interpretable directions in activation space corresponding to different agent perspectives. We further demonstrate a practical application: by instilling malicious agents into the LLM through internalized debate, then applying negative steering to suppress them, we show that distillation makes harmful behaviors easier to localize and control with smaller reductions in general performance compared to steering base models. Our findings offer a new perspective for understanding multi-agent capabilities in distilled models and provide practical guidelines for controlling internalized reasoning behaviors. Code available at https://github.com/johnsk95/latent_agents
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.