NTH

A Zeroth-Order Paradigm for LLM Preference Alignment

AuthorsPeter Chen, Xi Chen, Wotao Yin, Tianyi Lin

AffiliationsDepartment of Electrical Engineering and Computer Sciences (EECS) · University of California, Berkeley · Stern School of Business · New York University · Decision Intelligence Lab (Seattle) · DAMO Academy, Alibaba Group U.S · Department of Industrial Engineering and Operations Research (IEOR) · Columbia University

September 26, 2026 2 min read
Watch on YouTube
The one-line take

ComPO aligns language models using preference comparisons rather than direct likelihood optimization, aiming to reduce failures caused by small preference margins.

Key results

3
Preference margin threshold

Margin used to separate clean and low-margin noisy preference pairs.

100
Noisy pairs used

First noisy pairs used by ComPO in the Mistral-7B-Instruct experiment.

26.17%
Mistral AlpacaEval 2 LC

DPOclean plus ComPO length-controlled win rate on Mistral-7B-Instruct.

42.55%
Gemma-3 online AlpacaEval 2 LC

Length-controlled win rate after online damping and replay for Gemma-3-4B-it.

What the paper found

A Zeroth Order Paradigm for LLM Preference Alignment introduces Comparison-based Preference Optimization, or ComPO, to address likelihood displacement in direct alignment methods such as DPO and SimPO. Instead of differentiating a preference loss on low-margin pairs, ComPO perturbs a policy, uses a one-bit comparison oracle to check whether preferred-response likelihood rises while dispreferred-response likelihood falls, and aggregates signed perturbations into a sparse update. The practical method freezes most parameters and updates thresholded output-layer entries; preference pairs are split at margin threshold 3, with DPO or SimPO applied to clean pairs and ComPO applied to noisy ones. Its offline theory gives best-iterate convergence under smoothness, approximate gradient sparsity, and oracle compatibility, while online ComPO uses unlabeled generations to estimate a reverse-KL proxy for adaptive damping and replay. Across Mistral-7B, Llama-3, Gemma-2, Qwen3, and Gemma-3, ComPO improves AlpacaEval 2, Arena-Hard, and MT-Bench results. On Mistral-7B-Instruct, DPOclean plus ComPO reaches 26.17% length-controlled win rate on AlpacaEval 2, versus 23.89% for DPOclean, using only the first 100 noisy pairs. The method can update about 1.5M parameters, or 0.02% of a 7B model, and reports roughly 23 GB peak GPU memory for Llama-3-8B. Online damping plus replay raises Gemma-3-4B-it’s AlpacaEval 2 length-controlled win rate to 42.55%, compared with 40.00% for offline ComPO, with GPT-4.1-based evaluations used in the online experiments.

Original abstract

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we propose and analyze Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles. ComPO extracts directional information from these pairs without directly optimizing a differentiable preference loss on them. We establish a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective. We further introduce online ComPO, which retains the offline comparison mechanism and uses unlabeled policy generations for reverse-KL control relative to a reference policy. Following the coverage perspective of preference fine-tuning, we establish a performance guarantee for a basic constrained scheme under local coverage and in-distribution pairwise reward accuracy. Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 models demonstrate improvements over existing direct alignment methods, including length-controlled win rates, with pair-level diagnostics providing evidence consistent with mitigating likelihood displacement.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis