NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
AuthorsDelip Rao, Weiqiu You, Eric Wong, Chris Callison-Burch
This paper turns millions of NSF award abstracts into a new dataset for extracting scientific claims and research plans, aiming to power better claim verification and science analysis.
Key results
scientific claims in NSF-SciFy
manual error rate for fine-tuned claim extraction
manual error rate for fine-tuned investigation proposal extraction
What the paper found
NSF-SciFy introduces a new scientific-claim corpus mined from National Science Foundation award abstracts, using zero-shot prompting with Anthropic’s Claude-3.5-Sonnet to jointly extract factual claims and forward-looking investigation proposals. The full release contains 2.8 million claims from 400,000 NSF abstracts spanning all science and mathematics, with focused subsets NSF-SciFy-MATSCI at 114,000 claims from 16,042 materials-science awards and NSF-SciFy-20K at 135,000 claims from 20,001 awards across five NSF directorates. The paper shows that grant abstracts are stylistically distinct from their non-technical versions, and that the extracted content is high quality: manual evaluation found claim-extraction error at 2.6% and proposal-extraction error at 2.4%, while Claude-extracted claims had a slightly lower 2.1% error rate. To test utility, the authors fine-tune Mistral-7B-instruct-v0.3 and Qwen2.5-7B-Instruct with LoRA for three tasks: technical-to-non-technical abstract generation, claim extraction, and proposal extraction. Generation improves only modestly, with Mistral reaching BERTScore-F1 0.8561, but extraction tasks improve dramatically: Mistral reaches precision 0.7450, recall 0.7098, and F1 0.7097 for claims, and precision 0.7351, recall 0.7539, and F1 0.7261 for investigation proposals, with relative gains often exceeding 100%. The work positions NSF award abstracts as a novel source for large-scale scientific claim verification, discovery tracking, and meta-scientific analysis, and releases the datasets and models openly under Apache 2.0.
Original abstract
We introduce NSF-SciFy, a comprehensive dataset of scientific claims and investigation proposals extracted from National Science Foundation award abstracts. While previous scientific claim verification datasets have been limited in size and scope, NSF-SciFy represents a significant advance with 2.8 million claims from 400,000 abstracts spanning all science and mathematics disciplines. We present two focused subsets: NSF-SciFy-MatSci with 114,000 claims from materials science awards, and NSF-SciFy-20K with 135,000 claims across five NSF directorates. Using zero-shot prompting, we develop a scalable approach for joint extraction of scientific claims and investigation proposals. We demonstrate the dataset's utility through three downstream tasks: non-technical abstract generation, claim extraction, and investigation proposal extraction. Fine-tuning language models on our dataset yields substantial improvements, with relative gains often exceeding 100%, particularly for claim and proposal extraction tasks. Our error analysis reveals that extracted claims exhibit high precision but lower recall, suggesting opportunities for further methodological refinement. NSF-SciFy enables new research directions in large-scale claim verification, scientific discovery tracking, and meta-scientific analysis. Code and data are available at https://github.com/darpa-scify/NSFSciFy.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.