NTH

You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

AuthorsZiyang Luo, Zhongyao Chu, Xinjie He, Youting Wang, Xukui Qin, Runxiong Wu, Yan-Syuan Chen

August 23, 2026 3 min read
Watch on YouTube
The one-line take

YOPO lets a frozen language model reason, steer itself, and abstain when uncertain—all in one forward pass.

Key results

0.798
YOPO alphaNLI accuracy

One-pass three-way accuracy on Qwen2.5-1.5B, versus 0.375 for the frozen baseline.

0.888
Label-free transfer AUROC

HellaSwag transfer AUROC after reconstruction on Qwen2.5-1.5B, up from 0.836.

0.982
Supervised in-domain AUROC

Flagship one-pass gate in-domain AUROC on Qwen2.5-1.5B.

10/10
Cross-backbone win rate

One-pass YOPO beats the two-pass reference across ten backbones.

0.313
MuSiQue-F QA improvement

Answerable QA accuracy with chain-of-thought, up from 0.062.

What the paper found

YOPO, or You Only Pass Once, combines reasoning enhancement and abstention in one forward pass of a frozen language model. Tested primarily on Qwen2.5 at 1.5B, 3B, and 7B parameters, it uses a small conditional steering probe to write useful information into mid-layer residual states, while a fixed difference-of-means sufficiency direction detects when the context lacks enough evidence to answer. Because steering contaminates the state read by the abstention gate, YOPO trains a label-free reconstruction map with mean-squared error to recover the clean residual from steered activations, without sufficiency labels. On αNLI, one-pass three-way accuracy on Qwen2.5-1.5B rises from 0.375 for the frozen baseline to 0.798, exceeding the 0.753 two-pass reference and outperforming steering-only at 0.590 and gate-only at 0.560. The label-free correction improves 1.5B transfer AUROC on HellaSwag from 0.836 to 0.888, while supervised boosting reaches 0.982 in-domain but reduces transfer to 0.859, exposing a capacity–transfer trade-off. Across Qwen2.5, OLMo-2, TinyLlama, StableLM-2, SmolLM2, Phi-3-mini, and Mistral, the one-pass system beats the two-pass reference on 10/10 backbones. On native-label benchmarks including SQuAD2, RepLiQA, and MuSiQue, label-free sufficiency directions transfer better than trained gates. Finally, chain-of-thought raises answerable MuSiQue-F QA from 0.062 to 0.313 while the prefill-time gate remains protected from reasoning-induced contamination.

Original abstract

A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. This paper consolidates two research lines that address these on the same residual stream: a conditional steering probe writes the stream at mid-stack layers and recovers reasoning accuracy from a frozen backbone, and a zero-shot sufficiency direction reads the stream and abstains when information is insufficient. Deployed in one forward pass they interfere: the steering write shifts the state the direction reads, costing up to 8 AUROC points of cross-domain transfer on small models; a separate clean pass doubles inference cost. We keep the direction fixed and train a small network to reconstruct the pre-steering residual from the steered one -- mean-squared error on (steered, clean) pairs, no sufficiency labels -- and read the direction on the reconstruction. The resulting system, YOPO (You Only Pass Once), answers, steers, and abstains in one forward pass of a frozen Qwen2.5 backbone (1.5B/3B/7B). End to end, three-way accuracy more than doubles the frozen baseline (0.375->0.798 on 1.5B alphaNLI) and one pass beats the two-pass reference at every scale (0.798/0.830/0.893 vs 0.753/0.790/0.863) and on ten backbones across six model families. We chart the capacity-transfer frontier quantifying the principle that abstention should not be trained in; a source-side audit catches our own alphaNLI construction leaking a surface artifact, so architectural claims are anchored on native-label replications (SQuAD2, RepLiQA, MuSiQue); and on the standard four-domain suite we contribute, to our knowledge, the first answer-or-abstain benchmark, where our gate tops every in-domain dataset and the label-free direction is the only gate family to survive domain transfer.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis