Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
AuthorsJiacheng Miao, Jin Mu, Guanhua Chen, James Zou
Resources
Fisher-R1 trains LLM agents to choose statistically valid tests and draw more reliable scientific conclusions from real-world datasets.
Key results
Expert-audited open-ended hypothesis-testing tasks spanning economics, biology, and medicine.
Executable tasks used for supervised fine-tuning and reinforcement learning.
Verified Claude Sonnet 4.6 trajectories retained for supervised fine-tuning.
Single-trial strict success rate requiring a correct conclusion and numerically close p-value.
Average relative improvement in single-trial success reported by the paper.
What the paper found
Fisher-R1 targets a failure mode that ordinary coding-agent benchmarks miss: an LLM can execute valid code yet choose a statistically inappropriate test, report a misleading p-value, and draw a false discovery. The paper introduces P-Bench, an expert-audited benchmark of 425 open-ended hypothesis-testing tasks across economics, biology, and medicine; each task requires exploratory analysis, assumption checking, method selection, R execution, p-value reporting, and a reject-or-fail-to-reject decision. Fisher-R1 is an open-weight agent built from Qwen2.5-Coder backbones, first supervised-fine-tuned on 3,851 verified trajectories generated with Claude Sonnet 4.6, then optimized with DAPO reinforcement learning on 8,642 synthetic executable tasks. Its reward emphasizes p-value closeness in two-sided z-space, weighted 0.9, while conclusion consistency receives 0.1. Fisher-R1-14B reaches 33.0 percent Strict pass@1 on P-Bench’s hard split, compared with 26.3 percent for DeepSeek-V4-Pro and 30.5 percent for GPT-5.4, producing a 21 percent average relative improvement over DeepSeek-V4-Pro and gains up to 26 percent on the hardest tasks. The model also beats substantially larger open models, showing that verified statistical rewards and targeted reinforcement learning can matter more than parameter scaling. A central example contrasts GPT-5.4’s outlier-driven linear regression false positive with Fisher-R1’s appropriate Spearman test and correct failure to reject.
Original abstract
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchmarks fail to capture this failure mode, as they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data. We address this gap by building P-Bench, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine. Each task requires an agent to select a statistical method, compute a p-value, and draw a conclusion given only a scientific hypothesis and a dataset. We further introduce Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning. On P-Bench, Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks. Our results demonstrate that current LLM agents lack reliable statistical reasoning for hypothesis testing and that reinforcement learning on tasks with verified statistical reward substantially improves reliability.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.