ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
AuthorsHejia Geng, Zesen Huang, Haoyang Li, Wenbin Li, Koutian Wu, Zihan Zhou, Yuanbo Pang, Weihao Liu, Zigong Xu, Zhiping Li, Zongzheng Zhang, Chuanfei Dong, Jiankai Sun, Tianzhe Zheng, Fengyu Xie, Yue Ma, Yueheng Shi, Tong Xie, Zonglin Di, Xianrong Liu, Qucheng Gao, Yimin Liu, Jiaming Pan, Sheng Huang, Xiao-Han Ma, Lanqing Yuan, Zhenlin Zhu, Ziang Liu, Ziyang Xu, Junkai Wang, Kangkai Liang, Jiayi Xian, Zehong Zhao, Liuwei Xu, Jingxu Xie, Peijin Zhang, Qiang Gao, Chengyi Xing, Zhe Zhao, Xi Wang, Yaopeng Xing, Xing Meng, Zhenfei Yin, Yingcheng Wu, Ling Yang
AffiliationsSep 2026 Yuanbo Pang6 Weihao Liu7 Zigong Xu8 Zhiping Li5 Zongzheng Zhang1 Chuanfei Dong9 Jiankai Sun10 Tianzhe Zheng8 Fengyu Xie1 Yue Ma11 Yueheng Shi10 Junkai Wang1 Tong Xie12 Zonglin Di6,13 Xianrong
ScienceIDE turns scientific codebases into interactive training grounds where AI agents can learn to solve, verify, and improve scientific programming tasks.
Key results
ScienceIDE currently spans 64 environments across 27 scientific codebases.
The registry contains 2,812 executable scientific tasks.
Claude Fable 5.1 achieved 67.1% strict success on the 85-task benchmark.
Qwen3.5-9B improved from 0.240 to 0.576 after scientific-trajectory fine-tuning.
Held-out LAPS reward increased from 0.357 to 0.857 with verifier-guided reinforcement learning.
What the paper found
ScienceIDE converts scientific software repositories into agent-learnable environments by packaging pinned code, runtimes, official tests, expert-calibrated numerical checks, editable workspaces, and private verifiers. Its task factories generate repair, implementation, reproduction, acceleration, calibration, integration, and discovery challenges, but admit them only after execution proves that objectives are solvable, scientifically observable, and resistant to leakage. The current registry spans 64 environments across 27 codebases and contains 2,812 manufactured tasks, while a companion collection provides executable checks for numerical outputs and physical invariants. On the 85-task ScienceIDE-Hard benchmark with a one-hour budget, Anthropic’s Claude Fable 5.1 achieved 67.1% strict scientific success, ahead of OpenAI’s GPT-6 Astra at 63.1%; Claude Code, Codex, and Gemini CLI were used as execution harnesses, alongside models from Qwen and DeepSeek. Verified trajectories support supervised fine-tuning and reinforcement learning, including the PhAI-IDE-72B, PhAI-IDE-9B, and PhAI-IDE-4B family. Fine-tuning Qwen3.5-9B raised BBH Word Sorting from 0.240 to 0.576, suggesting transfer beyond scientific repair into reasoning and code benchmarks. For online learning, verifier-derived rewards and a truncation mask increased held-out LAPS reward from 0.357 to 0.857 and MITgcm-biogeo reward from 0.286 to 0.571, while preventing reinforcement learning from favoring prematurely shortened trajectories. The main limitation is coverage: most tasks concern reference-verifiable repair and implementation in computational physics and geoscience, not open-ended discovery or unseen codebases.
Original abstract
Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the scientific experience bottleneck. We introduce ScienceIDE, infrastructure for turning the world's scientific code into programmable environments for scientific agents. Guided by expert-defined scientific cases and acceptance criteria, agents transform repositories into executable environments that support task generation, execution, and scientific verification. These environments provide a shared foundation for supervised fine-tuning, reinforcement learning, and evaluation. Using verified interaction trajectories, we train PhAI-IDE-72B, PhAI-IDE-9B, and PhAI-IDE-4B. The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities. ScienceIDE lays the foundation for an integrated workspace for agent learning and scientific practice, making humanity's scientific software a shared substrate for developing scientific intelligence. Code: https://github.com/aitofound/ScienceIDE
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.