Codifying the Judge: Scalable Evaluation via Program Distillation
AuthorsTzu-Heng Huang, Shengqi Qiu, Frederic Sala
Resources
PAJAMA turns expensive, opaque LLM judges into fast, inspectable programs while escalating only uncertain cases to an LLM.
Key results
PAJAMA synthesizes 80 candidate Python judges before calibration and selection.
Average accuracy across 5 preference datasets.
Average evaluation throughput in samples per second.
Accuracy improvement achieved at 2.9× higher throughput than LLM-only evaluation.
Reward model trained on Prometheus examples labeled by PAJAMA.
PAJAMA labeling is estimated to be 50× cheaper than GPT-4 labeling for the Prometheus distillation setup.
What the paper found
Researchers at the University of Wisconsin–Madison introduce PAJAMA, a program-distillation framework that replaces repeated LLM-as-a-judge inference with a committee of transparent Python evaluators. Using Anthropic’s Claude Opus 4.6 to synthesize 80 candidate programs from 10 curated rubrics, PAJAMA calibrates scores with min–max normalization, selects reliable judges, aggregates their votes through weak supervision, and routes low-confidence cases to an LLM fallback. Across 5 preference datasets and four model families, the program-only system reaches 78.11% average accuracy—matching OLMo-2-13B-Instruct—while processing 439.41 samples per second. Its confidence router improves the OLMo-2-7B-Instruct baseline by +5.0% accuracy at 2.9× higher throughput. For reward modeling, Qwen2.5-3B-Instruct trained on PAJAMA labels achieves a 56.65 RewardBench average from Prometheus data, versus 54.32 from GPT-4 labels, while reducing labeling API cost by 50×. The key advantage is auditability: unlike opaque judges from systems such as OpenAI’s GPT-4, synthesized programs can be inspected, edited, and recalibrated; coding-agent calibration reduces average bias metrics from 39.78% to 35.58%.
Original abstract
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.