NTH
AI research

ACL-Verbatim: hallucination-free question answering for research

AuthorsGábor Recski, Szilveszter Tóth, Nadia Verdha, István Boros, Ádám Kovács

June 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper builds a benchmark and extractive QA system that answers research questions by copying exact text from ACL papers, aiming to reduce hallucinations in AI-assisted literature review.

Key results

120034
ACL papers processed

Entries from the ACL Anthology processed for corpus creation and indexing.

114475
PDFs converted to markdown

ACL Anthology PDFs converted to markdown with Docling.

100
Manual benchmark size

Gold benchmark of query–chunk pairs used for extraction evaluation.

47
Relevant chunks

Chunks in the gold benchmark annotated as relevant and containing gold evidence spans.

78
Gold evidence spans

Total annotated evidence spans in the manually labeled benchmark.

53.6
Best Word-F1

Best word-level F1 on the gold benchmark achieved by the 150M-parameter ModernBERT student.

What the paper found

ACL-Verbatim, from TU Wien and KR Labs, adapts the VerbatimRAG framework to the ACL Anthology to eliminate hallucinations in research-paper question answering by returning only verbatim text spans from retrieved documents. The authors process 120,034 ACL papers, convert 114,475 PDFs to markdown with Docling, and index section-aware chunks using full-text search plus dense retrieval with IBM’s granite-embedding-english-r2. To build supervision, they synthesize realistic queries with a ScIRGen-based pipeline, then manually annotate 100 query–chunk pairs using NLP researchers; 47 chunks are relevant and contain 78 gold evidence spans. On this benchmark, a 150M-parameter ModernBERT token classifier initialized from a reranker backbone achieves the best word-level F1 at 53.6, outperforming the strongest LLM extractor, GLM-5, at 48.7, while using 3–4 orders of magnitude fewer parameters. It also reaches 65.4 precision and abstains on most irrelevant chunks, which is crucial for high-precision evidence filtering in retrieval-augmented generation. The study argues that word-level overlap metrics are more appropriate than exact-span scores for this extraction task, and releases both the dataset and open-source pipeline as a blueprint for hallucination-free, explainable QA over scientific literature.

Original abstract

Academic researchers need efficient and reliable methods for collecting high-quality information from trusted sources, but modern tools for AI-assisted research still suffer from the tendency of Large Language Models (LLMs) to produce factually inaccurate or nonsensical output, commonly referred to as hallucinations. We apply the extractive question answering system VerbatimRAG to research papers in the ACL Anthology, directly mapping user queries to verbatim text spans in retrieved documents. We contribute a novel ground truth dataset for the task of mapping user queries to relevant text spans in research papers, and use it to train and evaluate a variety of extractive models. Human annotation is performed by NLP researchers and is based on synthetic user queries generated using a custom pipeline based on the ScIRGen methodology, paired with chunks of research papers retrieved by VerbatimRAG. On this benchmark, a 150M-parameter ModernBERT token classifier trained on silver supervision from our pipeline achieves the best word-level F1 (53.6), ahead of the strongest evaluated LLM extractor (48.7).

Read the original paper