Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models
AuthorsFei Yuan, Changjiang Gao, Yilei Tu, Yifeng Liu, Shujian Huang, Yu Qiao
This work argues that successful fine-tuning depends less on how far a model moves than on choosing the right direction while keeping behavioral drift within a fixed budget.
Key results
SmolInstruct examples used for QA-only scientific reasoning fine-tuning
Lego-MT translation pairs used for multilingual adaptation
The translation experiments cover more than 100 languages
Aggregate xCOMET score for the strongest split Layer-Selective Tuning configuration
Anchored KL divergence from the Qwen3-14B reference model
What the paper found
This paper reframes fine-tuning as direction selection under a fixed behavioral-drift budget. Using anchored KL divergence, it models local behavior with a Fisher-information geometry in which drift measures distance from the reference model, while the update direction determines task improvement per unit drift; the ideal local direction is the natural gradient, −F⁻¹g. The idea is tested on Qwen3-8B and Qwen3-14B in a stringent QA-only setting: models train on answers without reasoning trajectories, yet must preserve multi-step reasoning at inference. Full fine-tuning and KL regularization often trade target gains for capability loss, whereas Layer-Selective Tuning, especially split configurations such as b4t16, exposes more efficient directions. Training uses 300K SmolInstruct examples for scientific reasoning and approximately 2.8M Lego-MT translation pairs covering more than 100 languages, evaluated with FLORES-101 and xCOMET. On Qwen3-14B, b4t16 reaches 58.31 xCOMET for multilingual translation with only 0.14 KL drift, outperforming the 55.21 reference score. The same directional structure also improves initialization for subsequent reinforcement learning. The central claim is that, once behavioral change is constrained, fine-tuning quality depends less on how far parameters move than on where that movement occurs.
Original abstract
Fine-tuning instruct models often improves target performance while inducing behavioral drift from the reference model, which can degrade existing capabilities. Rather than treating this drift as an uncontrolled consequence of optimization, we specify a behavioral drift budget before optimization and ask how to boost the target-task performance within it. Locally, behavioral drift induces a shared geometry anchored at the reference model, with the drift budget defining a boundary within this space. In this space, drift determines distance from the reference, leaving update direction as the remaining degree of freedom. Fine-tuning updates can therefore be compared through their directional efficiency, naturally reformulating fine-tuning as a direction-selection problem. This reformulation makes a concrete prediction: changing the accessible directions can qualitatively alter the outcome of fine-tuning. We test this prediction in a stringent QA-only setting, where strong instruct models are fine-tuned only on final answers but must still generate multi-step reasoning at inference. Despite this mismatch, a coarse layer-selective probe reverses the failure of QA-only fine-tuning and reveals the existence of effective directions, with multiple neighboring configurations improving target performance while preserving reasoning and general capabilities. Across Qwen3-8B and Qwen3-14B, these directions substantially improve scientific reasoning and multilingual translation. Over more than 100 languages, the resulting models match or outperform dedicated translation systems and provide a stronger initialization for subsequent reinforcement learning. Our results suggest that fine-tuning is not just about how much a model changes, but how that change is spent. https://github.com/CONE-MT/DCO and https://huggingface.co/collections/LLaMAX/dco
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.