A Zeroth-Order Paradigm for LLM Preference Alignment
AuthorsPeter Chen, Xi Chen, Wotao Yin, Tianyi Lin
AffiliationsDepartment of Electrical Engineering and Computer Sciences (EECS) · University of California, Berkeley · Stern School of Business · New York University · Decision Intelligence Lab (Seattle) · DAMO Academy, Alibaba Group U.S · Department of Industrial Engineering and Operations Research (IEOR) · Columbia University
Resources
ComPO aligns language models using preference comparisons rather than direct likelihood optimization, aiming to reduce failures caused by small preference margins.
Key results
Margin used to separate clean and low-margin noisy preference pairs.
First noisy pairs used by ComPO in the Mistral-7B-Instruct experiment.
DPOclean plus ComPO length-controlled win rate on Mistral-7B-Instruct.
Length-controlled win rate after online damping and replay for Gemma-3-4B-it.
What the paper found
A Zeroth Order Paradigm for LLM Preference Alignment introduces Comparison-based Preference Optimization, or ComPO, to address likelihood displacement in direct alignment methods such as DPO and SimPO. Instead of differentiating a preference loss on low-margin pairs, ComPO perturbs a policy, uses a one-bit comparison oracle to check whether preferred-response likelihood rises while dispreferred-response likelihood falls, and aggregates signed perturbations into a sparse update. The practical method freezes most parameters and updates thresholded output-layer entries; preference pairs are split at margin threshold 3, with DPO or SimPO applied to clean pairs and ComPO applied to noisy ones. Its offline theory gives best-iterate convergence under smoothness, approximate gradient sparsity, and oracle compatibility, while online ComPO uses unlabeled generations to estimate a reverse-KL proxy for adaptive damping and replay. Across Mistral-7B, Llama-3, Gemma-2, Qwen3, and Gemma-3, ComPO improves AlpacaEval 2, Arena-Hard, and MT-Bench results. On Mistral-7B-Instruct, DPOclean plus ComPO reaches 26.17% length-controlled win rate on AlpacaEval 2, versus 23.89% for DPOclean, using only the first 100 noisy pairs. The method can update about 1.5M parameters, or 0.02% of a 7B model, and reports roughly 23 GB peak GPU memory for Llama-3-8B. Online damping plus replay raises Gemma-3-4B-it’s AlpacaEval 2 length-controlled win rate to 42.55%, compared with 40.00% for offline ComPO, with GPT-4.1-based evaluations used in the online experiments.
Original abstract
Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we propose and analyze Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles. ComPO extracts directional information from these pairs without directly optimizing a differentiable preference loss on them. We establish a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective. We further introduce online ComPO, which retains the offline comparison mechanism and uses unlabeled policy generations for reverse-KL control relative to a reference policy. Following the coverage perspective of preference fine-tuning, we establish a performance guarantee for a basic constrained scheme under local coverage and in-distribution pairwise reward accuracy. Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 models demonstrate improvements over existing direct alignment methods, including length-controlled win rates, with pair-level diagnostics providing evidence consistent with mitigating likelihood displacement.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.