On the Limits of LLM-as-Judge for Scientific Novelty Assessment
AuthorsSoumitra Sinhahajari, Navonil Majumder, Soujanya Poria
Resources
This paper shows that LLM judges can be badly overconfident about scientific novelty, and introduces a benchmark revealing that they often call model-generated research questions more novel than domain experts do.
Key results
Total research questions in the benchmark
Links between research questions and cited papers
Extracted gap statements used to form the benchmark
Agreement between humans and LLM judges on 50 sampled items
Two-judge combined win rate for non-obviousness in comparative scoring
What the paper found
This paper argues that using LLM-as-judge for scientific novelty assessment creates a “novelty mirage,” so it builds RQ-Bench, a benchmark of 1,434 research questions reconstructed from 746 recent arXiv computer science papers, with 2,464 RQ-citation links and 3,151 grounded gap statements. The authors evaluate frontier models including gpt-5.5, gemini-3.1-pro, and deepseek-v4-pro on three novelty dimensions—originality, gap addressing, and non-obviousness—under standalone and comparative judging. LLM judges consistently assign high novelty to model-generated questions, but the effect strengthens in comparative evaluation: for gpt-5.5, the two-judge combined win rate rises from 27.2% to 49.1% on non-obviousness when switching from standalone to comparative scoring. Human experts on 50 cs.CL and cs.LG samples disagree sharply, preferring author-anchored questions, with expert-judge agreement falling to 22% on non-obviousness. The paper also shows that generated questions are often narrow and source-bound; when narrowness is scored explicitly, it aligns with expert judgments better than semantic similarity, and updating the generation prompt reduces source-boundedness from 1.00 to 0.47 for gemini-3.1-pro while also lowering its novelty scores. The main conclusion is that current LLM judges are not reliable for scientific novelty because they conflate polished gap language with genuine non-obviousness and broader scope.
Original abstract
LLMs are increasingly used to generate and judge scientific ideas. This makes novelty evaluation a central problem. Full idea evaluation is difficult because it often requires judging a method, its feasibility, and its empirical promise. We therefore study a cleaner upstream object: the research question (RQ). RQ generation is a prerequisite for scientific ideation, and RQs can be compared against questions pursued in real papers. We introduce RQ-Bench, a benchmark built from recent arXiv papers. For each paper, we reconstruct author-anchored RQs from its cited background, gaps, and contributions. These RQs are not the only valid questions for the same background. They are author-anchored reference points for testing novelty judgments. We evaluate model-generated RQs with standalone LLM judging, comparative LLM judging, and human expert evaluation. LLM judges consistently rate model-generated RQs as highly novel, producing a novelty mirage; in comparative evaluations, this preference becomes even stronger. Domain experts, however, reach the opposite conclusion and prefer the author-anchored reference questions. We further find that many generated RQs are narrow or source-bound, a dimension that LLM judges often miss unless explicitly tested. Overall, the contradictory novelty evaluations between LLM judges and human experts raise a serious concern about the reliability of using LLMs to assess the scientific novelty of research questions.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.