NTH

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher

AuthorsArda Uzunoglu, Alvin Zhang, Daniel Khashabi

June 10, 2026 2 min read
Watch on YouTube
The one-line take

This paper learns when to trust weak labels so a stronger model can turn imperfect supervision into nearly lossless training and even bootstrap itself repeatedly.

Key results

87.1
World knowledge NTF accuracy

Qwen3-0.6B teacher to Qwen3-14B student, average world knowledge result

27.9
Quantitative reasoning NTF accuracy

Qwen3-4B student under Qwen3-1.7B teacher

44.1
Strategy games NTF accuracy

Qwen3-14B student under Qwen3-0.6B teacher

0.96
NTF AUC

Best trust-function discrimination reported for chess using Qwen3-0.6B

0.98
NTF purity

Highest retained-label purity reported for world knowledge trust filtering

48.2
Chaining result

Final Qwen3-14B accuracy in the weak-to-strong chain on strategy games

What the paper found

This Johns Hopkins University paper introduces neural trust functions, a data-selection method for weak-to-strong generalization that scores each weak label using the weak teacher’s internal activations rather than output confidence. Trained on a labeled source set and deployed zero-shot under in-domain distribution shift, the trust function filters weak supervision before training a stronger student. Across world knowledge, quantitative reasoning, and strategy games, the method achieves near-lossless recovery relative to ground-truth training and sometimes surpasses it: in world knowledge, the Qwen3-0.6B teacher yields 61.6% to 87.1% student accuracy and matches or beats gold-label training in 5 of 8 settings; in quantitative reasoning, it reaches 22.0 on Qwen3-4B and 27.9 on Qwen3-8B, close to or above ground truth; in chess, it reaches 44.1 on Qwen3-14B versus 39.9 for ground truth. The trust model itself is well calibrated, with AUC up to 0.96 and purity up to 0.98, and chaining the procedure across generations compounds gains, culminating in 48.2 on Qwen3-14B, ahead of the 40.0 ground-truth baseline. Mechanistically, the selected data are easier, more label-correct, and induce more coherent low-rank gradients, while some apparent false positives are actually stronger alternatives than the dataset labels. A risk-controlled Hoeffding calibration rule further turns trust scores into a practical thresholding mechanism, retaining 16.1% of the deployment pool at a target noise rate of 0.1.

Original abstract

Weak-to-strong generalization studies how to improve a strong student using supervision from a weaker teacher when reliable labels are scarce. We view this primarily as a data selection problem, where the key challenge is to identify which weak labels are reliable enough to serve as a training signal. To address this, we introduce trust functions that assign each weak label a scalar trust score and use these scores to filter weak supervision. Across several domains, including world knowledge, quantitative reasoning, and strategy games, trust filtering yields students that match and sometimes surpass ground-truth supervision, achieving near-lossless weak-to-strong generalization. Moreover, trust functions enable an iterative weak-to-strong chain that compounds gains by training a student and reusing it as the next teacher, amplifying the gains. There are several mechanisms to which advantage of trust functions can be attributed.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis