You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model
AuthorsZiyang Luo, Zhongyao Chu, Xinjie He, Youting Wang, Xukui Qin, Runxiong Wu, Yan-Syuan Chen
Resources
YOPO lets a frozen language model reason, steer itself, and abstain when uncertain—all in one forward pass.
Key results
One-pass three-way accuracy on Qwen2.5-1.5B, versus 0.375 for the frozen baseline.
HellaSwag transfer AUROC after reconstruction on Qwen2.5-1.5B, up from 0.836.
Flagship one-pass gate in-domain AUROC on Qwen2.5-1.5B.
One-pass YOPO beats the two-pass reference across ten backbones.
Answerable QA accuracy with chain-of-thought, up from 0.062.
What the paper found
YOPO, or You Only Pass Once, combines reasoning enhancement and abstention in one forward pass of a frozen language model. Tested primarily on Qwen2.5 at 1.5B, 3B, and 7B parameters, it uses a small conditional steering probe to write useful information into mid-layer residual states, while a fixed difference-of-means sufficiency direction detects when the context lacks enough evidence to answer. Because steering contaminates the state read by the abstention gate, YOPO trains a label-free reconstruction map with mean-squared error to recover the clean residual from steered activations, without sufficiency labels. On αNLI, one-pass three-way accuracy on Qwen2.5-1.5B rises from 0.375 for the frozen baseline to 0.798, exceeding the 0.753 two-pass reference and outperforming steering-only at 0.590 and gate-only at 0.560. The label-free correction improves 1.5B transfer AUROC on HellaSwag from 0.836 to 0.888, while supervised boosting reaches 0.982 in-domain but reduces transfer to 0.859, exposing a capacity–transfer trade-off. Across Qwen2.5, OLMo-2, TinyLlama, StableLM-2, SmolLM2, Phi-3-mini, and Mistral, the one-pass system beats the two-pass reference on 10/10 backbones. On native-label benchmarks including SQuAD2, RepLiQA, and MuSiQue, label-free sufficiency directions transfer better than trained gates. Finally, chain-of-thought raises answerable MuSiQue-F QA from 0.062 to 0.313 while the prefill-time gate remains protected from reasoning-induced contamination.
Original abstract
A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. This paper consolidates two research lines that address these on the same residual stream: a conditional steering probe writes the stream at mid-stack layers and recovers reasoning accuracy from a frozen backbone, and a zero-shot sufficiency direction reads the stream and abstains when information is insufficient. Deployed in one forward pass they interfere: the steering write shifts the state the direction reads, costing up to 8 AUROC points of cross-domain transfer on small models; a separate clean pass doubles inference cost. We keep the direction fixed and train a small network to reconstruct the pre-steering residual from the steered one -- mean-squared error on (steered, clean) pairs, no sufficiency labels -- and read the direction on the reconstruction. The resulting system, YOPO (You Only Pass Once), answers, steers, and abstains in one forward pass of a frozen Qwen2.5 backbone (1.5B/3B/7B). End to end, three-way accuracy more than doubles the frozen baseline (0.375->0.798 on 1.5B alphaNLI) and one pass beats the two-pass reference at every scale (0.798/0.830/0.893 vs 0.753/0.790/0.863) and on ten backbones across six model families. We chart the capacity-transfer frontier quantifying the principle that abstention should not be trained in; a source-side audit catches our own alphaNLI construction leaking a surface artifact, so architectural claims are anchored on native-label replications (SQuAD2, RepLiQA, MuSiQue); and on the standard four-domain suite we contribute, to our knowledge, the first answer-or-abstain benchmark, where our gate tops every in-domain dataset and the label-free direction is the only gate family to survive domain transfer.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.