Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
AuthorsArda Uzunoglu, Alvin Zhang, Daniel Khashabi
Resources
This paper learns when to trust weak labels so a stronger model can turn imperfect supervision into nearly lossless training and even bootstrap itself repeatedly.
Key results
Qwen3-0.6B teacher to Qwen3-14B student, average world knowledge result
Qwen3-4B student under Qwen3-1.7B teacher
Qwen3-14B student under Qwen3-0.6B teacher
Best trust-function discrimination reported for chess using Qwen3-0.6B
Highest retained-label purity reported for world knowledge trust filtering
Final Qwen3-14B accuracy in the weak-to-strong chain on strategy games
What the paper found
This Johns Hopkins University paper introduces neural trust functions, a data-selection method for weak-to-strong generalization that scores each weak label using the weak teacher’s internal activations rather than output confidence. Trained on a labeled source set and deployed zero-shot under in-domain distribution shift, the trust function filters weak supervision before training a stronger student. Across world knowledge, quantitative reasoning, and strategy games, the method achieves near-lossless recovery relative to ground-truth training and sometimes surpasses it: in world knowledge, the Qwen3-0.6B teacher yields 61.6% to 87.1% student accuracy and matches or beats gold-label training in 5 of 8 settings; in quantitative reasoning, it reaches 22.0 on Qwen3-4B and 27.9 on Qwen3-8B, close to or above ground truth; in chess, it reaches 44.1 on Qwen3-14B versus 39.9 for ground truth. The trust model itself is well calibrated, with AUC up to 0.96 and purity up to 0.98, and chaining the procedure across generations compounds gains, culminating in 48.2 on Qwen3-14B, ahead of the 40.0 ground-truth baseline. Mechanistically, the selected data are easier, more label-correct, and induce more coherent low-rank gradients, while some apparent false positives are actually stronger alternatives than the dataset labels. A risk-controlled Hoeffding calibration rule further turns trust scores into a practical thresholding mechanism, retaining 16.1% of the deployment pool at a target noise rate of 0.1.
Original abstract
Weak-to-strong generalization studies how to improve a strong student using supervision from a weaker teacher when reliable labels are scarce. We view this primarily as a data selection problem, where the key challenge is to identify which weak labels are reliable enough to serve as a training signal. To address this, we introduce trust functions that assign each weak label a scalar trust score and use these scores to filter weak supervision. Across several domains, including world knowledge, quantitative reasoning, and strategy games, trust filtering yields students that match and sometimes surpass ground-truth supervision, achieving near-lossless weak-to-strong generalization. Moreover, trust functions enable an iterative weak-to-strong chain that compounds gains by training a student and reusing it as the next teacher, amplifying the gains. There are several mechanisms to which advantage of trust functions can be attributed.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.