NTH

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

AuthorsLiwei Dong, Jiahao Zhao, Nan Xu

July 29, 2026 2 min read
Watch on YouTube
The one-line take

SciConsolidate turns successful and failed coding experiences into reusable scientific procedures that help smaller language models improve at computational problem solving.

Key results

3.85
Qwen3.6-27B runtime sub-step gain

Aggregate SciCode improvement in sub-step accuracy from runtime procedure injection.

6.26
Qwen3.6-27B runtime main-problem gain

Aggregate SciCode improvement in main-problem accuracy from runtime procedure injection.

2,095
Procedure-guided SFT examples

Retained training examples after generation and executable filtering.

1,985
No-procedure SFT examples

Retained matched-control training examples after executable filtering.

5.62
9B guided-over-base sub-step gain

Aggregate procedure-free improvement of the guided 9B student over the original Qwen3.5-9B model.

11.25
9B guided-over-base main-problem gain

Aggregate procedure-free improvement of the guided 9B student over the original Qwen3.5-9B model.

What the paper found

Researchers at ScienceOne AI and Wenge AI introduce SciConsolidate, a pipeline that turns verified scientific-computing execution traces into persistent model capability. On SciCode, the system contrasts successful and failed rollouts, groups recurring errors such as interface misuse, array-shape mistakes, and dependency-flow breaks, synthesizes family-level procedures, validates them on a development split, and generates answer-free scientific queries. Because the smaller target may not operationalize abstract procedures, the stronger Qwen3.6-27B concretizer converts them into executable code supervision, while Qwen3.5-9B learns through standard procedure-free supervised fine-tuning. Runtime procedure injection raises Qwen3.6-27B by 3.85 sub-step points and 6.26 main-problem points, but gives Qwen3.5-9B essentially no aggregate main-problem improvement, demonstrating an abstraction–execution gap. The guided training pool contains 2,095 retained examples, compared with 1,985 for the matched no-procedure control. After training, the 9B procedure-guided student improves over the original 9B model by 5.62 sub-step points and 11.25 main-problem points under procedure-free deployment, and it exceeds the no-procedure SFT control by 3.89 and 6.25 points respectively. Failure-family diagnostics show the strongest transfer for simulation, array or numerical reasoning, and logic or dependency flow, while broad numerical-method reasoning and specification-sensitive interfaces remain weak. GLM-5.1 from Z.ai is used for failure labeling, and the authors present the result as one evaluated experience-to-capability pass rather than autonomous multi-round self-improvement.

Original abstract

Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather than cross-task computational mechanisms; and a weaker target model may be unable to operationalize an otherwise valid abstract procedure - an abstraction-execution gap. We introduce SciConsolidate, which contrasts verified successes and failures to induce cross-task procedures, selects them through a development-validation gate, and uses failure-informed, answer-free query synthesis to expand the consolidation data without requiring pre-existing reference answers. Because the target model may not directly execute these abstractions, a stronger model concretizes them into executable code supervision for standard, procedure-free SFT; a matched no-procedure teacher branch isolates the value of procedural guidance. On SciCode, runtime procedure injection improves Qwen3.6-27B by +3.85/+6.26 sub-step/main-problem points, but yields almost no aggregate main-problem gain for Qwen3.5-9B, providing operational evidence of the abstraction-execution gap. After procedure-guided concretization, the 9B student improves under procedure-free deployment by +3.89/+6.25 points over the no-procedure SFT control and by +5.62/+11.25 over the original 9B model. These results establish an experience-to-capability pathway for scientific computing and provide a practical starting point for scaling self-improving scientific assistance.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis