NTH

On the Limits of LLM-as-Judge for Scientific Novelty Assessment

AuthorsSoumitra Sinhahajari, Navonil Majumder, Soujanya Poria

July 5, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that LLM judges can be badly overconfident about scientific novelty, and introduces a benchmark revealing that they often call model-generated research questions more novel than domain experts do.

Key results

1434
RQ-Bench size

Total research questions in the benchmark

2464
RQ-citation links

Links between research questions and cited papers

3151
Grounded gap statements

Extracted gap statements used to form the benchmark

22%
Expert agreement on non-obviousness

Agreement between humans and LLM judges on 50 sampled items

49.1%
gpt-5.5 comparative combined win rate

Two-judge combined win rate for non-obviousness in comparative scoring

What the paper found

This paper argues that using LLM-as-judge for scientific novelty assessment creates a “novelty mirage,” so it builds RQ-Bench, a benchmark of 1,434 research questions reconstructed from 746 recent arXiv computer science papers, with 2,464 RQ-citation links and 3,151 grounded gap statements. The authors evaluate frontier models including gpt-5.5, gemini-3.1-pro, and deepseek-v4-pro on three novelty dimensions—originality, gap addressing, and non-obviousness—under standalone and comparative judging. LLM judges consistently assign high novelty to model-generated questions, but the effect strengthens in comparative evaluation: for gpt-5.5, the two-judge combined win rate rises from 27.2% to 49.1% on non-obviousness when switching from standalone to comparative scoring. Human experts on 50 cs.CL and cs.LG samples disagree sharply, preferring author-anchored questions, with expert-judge agreement falling to 22% on non-obviousness. The paper also shows that generated questions are often narrow and source-bound; when narrowness is scored explicitly, it aligns with expert judgments better than semantic similarity, and updating the generation prompt reduces source-boundedness from 1.00 to 0.47 for gemini-3.1-pro while also lowering its novelty scores. The main conclusion is that current LLM judges are not reliable for scientific novelty because they conflate polished gap language with genuine non-obviousness and broader scope.

Original abstract

LLMs are increasingly used to generate and judge scientific ideas. This makes novelty evaluation a central problem. Full idea evaluation is difficult because it often requires judging a method, its feasibility, and its empirical promise. We therefore study a cleaner upstream object: the research question (RQ). RQ generation is a prerequisite for scientific ideation, and RQs can be compared against questions pursued in real papers. We introduce RQ-Bench, a benchmark built from recent arXiv papers. For each paper, we reconstruct author-anchored RQs from its cited background, gaps, and contributions. These RQs are not the only valid questions for the same background. They are author-anchored reference points for testing novelty judgments. We evaluate model-generated RQs with standalone LLM judging, comparative LLM judging, and human expert evaluation. LLM judges consistently rate model-generated RQs as highly novel, producing a novelty mirage; in comparative evaluations, this preference becomes even stronger. Domain experts, however, reach the opposite conclusion and prefer the author-anchored reference questions. We further find that many generated RQs are narrow or source-bound, a dimension that LLM judges often miss unless explicitly tested. Overall, the contradictory novelty evaluations between LLM judges and human experts raise a serious concern about the reliability of using LLMs to assess the scientific novelty of research questions.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis