NTH

Codifying the Judge: Scalable Evaluation via Program Distillation

AuthorsTzu-Heng Huang, Shengqi Qiu, Frederic Sala

July 30, 2026 2 min read
Watch on YouTube
The one-line take

PAJAMA turns expensive, opaque LLM judges into fast, inspectable programs while escalating only uncertain cases to an LLM.

Key results

80
Candidate programs synthesized

PAJAMA synthesizes 80 candidate Python judges before calibration and selection.

78.11%
Average programmatic-judge accuracy

Average accuracy across 5 preference datasets.

439.41
Programmatic-judge throughput

Average evaluation throughput in samples per second.

+5.0%
OLMo-2-7B-Instruct routing gain

Accuracy improvement achieved at 2.9× higher throughput than LLM-only evaluation.

56.65
RewardBench average with PAJAMA labels

Reward model trained on Prometheus examples labeled by PAJAMA.

50×
API cost reduction

PAJAMA labeling is estimated to be 50× cheaper than GPT-4 labeling for the Prometheus distillation setup.

What the paper found

Researchers at the University of Wisconsin–Madison introduce PAJAMA, a program-distillation framework that replaces repeated LLM-as-a-judge inference with a committee of transparent Python evaluators. Using Anthropic’s Claude Opus 4.6 to synthesize 80 candidate programs from 10 curated rubrics, PAJAMA calibrates scores with min–max normalization, selects reliable judges, aggregates their votes through weak supervision, and routes low-confidence cases to an LLM fallback. Across 5 preference datasets and four model families, the program-only system reaches 78.11% average accuracy—matching OLMo-2-13B-Instruct—while processing 439.41 samples per second. Its confidence router improves the OLMo-2-7B-Instruct baseline by +5.0% accuracy at 2.9× higher throughput. For reward modeling, Qwen2.5-3B-Instruct trained on PAJAMA labels achieves a 56.65 RewardBench average from Prometheus data, versus 54.32 from GPT-4 labels, while reducing labeling API cost by 50×. The key advantage is auditability: unlike opaque judges from systems such as OpenAI’s GPT-4, synthesized programs can be inspected, edited, and recalibrated; coding-agent calibration reduces average bias metrics from 39.78% to 35.58%.

Original abstract

LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis