From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis
AuthorsLiwei Dong, Jiahao Zhao, Nan Xu
Resources
SciConsolidate turns successful and failed coding experiences into reusable scientific procedures that help smaller language models improve at computational problem solving.
Key results
Aggregate SciCode improvement in sub-step accuracy from runtime procedure injection.
Aggregate SciCode improvement in main-problem accuracy from runtime procedure injection.
Retained training examples after generation and executable filtering.
Retained matched-control training examples after executable filtering.
Aggregate procedure-free improvement of the guided 9B student over the original Qwen3.5-9B model.
Aggregate procedure-free improvement of the guided 9B student over the original Qwen3.5-9B model.
What the paper found
Researchers at ScienceOne AI and Wenge AI introduce SciConsolidate, a pipeline that turns verified scientific-computing execution traces into persistent model capability. On SciCode, the system contrasts successful and failed rollouts, groups recurring errors such as interface misuse, array-shape mistakes, and dependency-flow breaks, synthesizes family-level procedures, validates them on a development split, and generates answer-free scientific queries. Because the smaller target may not operationalize abstract procedures, the stronger Qwen3.6-27B concretizer converts them into executable code supervision, while Qwen3.5-9B learns through standard procedure-free supervised fine-tuning. Runtime procedure injection raises Qwen3.6-27B by 3.85 sub-step points and 6.26 main-problem points, but gives Qwen3.5-9B essentially no aggregate main-problem improvement, demonstrating an abstraction–execution gap. The guided training pool contains 2,095 retained examples, compared with 1,985 for the matched no-procedure control. After training, the 9B procedure-guided student improves over the original 9B model by 5.62 sub-step points and 11.25 main-problem points under procedure-free deployment, and it exceeds the no-procedure SFT control by 3.89 and 6.25 points respectively. Failure-family diagnostics show the strongest transfer for simulation, array or numerical reasoning, and logic or dependency flow, while broad numerical-method reasoning and specification-sensitive interfaces remain weak. GLM-5.1 from Z.ai is used for failure labeling, and the authors present the result as one evaluated experience-to-capability pass rather than autonomous multi-round self-improvement.
Original abstract
Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather than cross-task computational mechanisms; and a weaker target model may be unable to operationalize an otherwise valid abstract procedure - an abstraction-execution gap. We introduce SciConsolidate, which contrasts verified successes and failures to induce cross-task procedures, selects them through a development-validation gate, and uses failure-informed, answer-free query synthesis to expand the consolidation data without requiring pre-existing reference answers. Because the target model may not directly execute these abstractions, a stronger model concretizes them into executable code supervision for standard, procedure-free SFT; a matched no-procedure teacher branch isolates the value of procedural guidance. On SciCode, runtime procedure injection improves Qwen3.6-27B by +3.85/+6.26 sub-step/main-problem points, but yields almost no aggregate main-problem gain for Qwen3.5-9B, providing operational evidence of the abstraction-execution gap. After procedure-guided concretization, the 9B student improves under procedure-free deployment by +3.89/+6.25 points over the no-procedure SFT control and by +5.62/+11.25 over the original 9B model. These results establish an experience-to-capability pathway for scientific computing and provide a practical starting point for scaling self-improving scientific assistance.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.