JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
AuthorsYubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman
AffiliationsCarnegie Mellon University · {yubol, yidim, rk2x, rpadman}@andrew.cmu.edu
Resources
A cheap judge handles confident cases while uncertain ones are escalated, preserving nearly all of a powerful LLM judge's accuracy for a fraction of the cost.
Key results
Total base judgments across RewardBench, JudgeBench, HaluEval, and final-answer tasks.
Estimated US dollars per 1,000 judgments.
Median response latency in seconds on the matched timing panel.
JEV accuracy on difficult correctness judgments, versus 93.1% for GPT-6.
Accuracy-point improvement of the frozen JEV-to-GPT-6 cascade over GPT-6 alone on 1,610 held-out pairs.
Share of 570 new workload pairs escalated while matching GPT-6 accuracy.
What the paper found
JEV-as-a-Judge evaluates responses with a typed, decision-only interface: instead of generating a rationale, JEV returns label probabilities, and the highest probability becomes confidence. The proposed cascade accepts JEV when confidence exceeds a locally chosen threshold and escalates uncertain cases to a reasoning judge such as OpenAI’s GPT-6 Astra. Across 5,172 judgments spanning RewardBench, JudgeBench, HaluEval, and final-answer adjudication, JEV costs $0.044 per 1,000 judgments and has 0.15-second median latency, while matching GPT-6 within roughly three points when verdicts can be read directly from text, including 92.5% on RewardBench and 87.3% versus 88.4% on HaluEval. Its weakness is derived correctness: on JudgeBench, JEV scores 78.6% versus GPT-6’s 93.1%, with especially large deficits on math, code, and logic; style-adversarial RM-Bench pairs also induce confident errors. Confidence nevertheless concentrates most errors in the escalation region. With a threshold fixed before evaluation, the JEV-to-GPT-6 cascade on 1,610 held-out pairs is 0.9 points more accurate than GPT-6 alone while costing 41% of its fee. In a pre-specified live test on 570 new correctness pairs, workload-specific thresholds match GPT-6’s accuracy exactly, escalating 74% of cases. Comparisons with Claude Sonnet 5, Gemini 3 Flash, and Qwen3 models show that the advantage comes from calibrated routing, not merely a typed output format. The paper’s operational conclusion is to validate thresholds locally, average presentation orders when relevant, and avoid relying on confidence for reference-free prose or adversarial style judgments.
Original abstract
LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.