NTH

OpenAI4S: Code as Action, Science as Sessions

AuthorsGongbo Zhang, Hao Li, Yu Wang, Mujie Lin, Liuzhenghao Lv, Yicheng Mao, Yimi Wang, Jun Zhu, Minhan Tang, Zhengxiang Jiang, Yusong Wang, Jiayu Yao, Kunpeng Ning, Dawei Pang, Yonghong Tian, OpenAI4S Community, Yuyang Liu, Li Yuan

AffiliationsPeking University Shenzhen Graduate School, Shenzhen, 518055, China · These authors contributed equally to this work · Tsinghua-Peking Joint Center for Life Sciences, School of Life Sciences, Tsinghua University, Beijing, China · Independent Researcher · Beijing Yuankong Intelligent Technology Co., Ltd., Beijing, China

September 22, 2026 2 min read
Watch on YouTube
The one-line take

OpenAI4S turns AI-driven scientific research into inspectable, resumable sessions by treating executable code as the agent’s actions and tracking every computational step.

Key results

36
Evaluation scenarios

Research scenarios spanning six scientific tasks.

7.83
OpenAI4S macro score

Overall repository-level score across the six-task evaluation.

6.36
GLM-5.2 baseline score

Score for GLM-5.2 operating through the Claude Code harness.

5.79
Kimi-K3 baseline score

Score for Kimi-K3 operating through the Claude Code harness.

604
Bundled scientific Skills

Runnable scientific recipes and helper modules included with OpenAI4S.

What the paper found

OpenAI4S is an open-source scientific research agent built around “Code as Action, Science as Sessions.” Instead of treating each tool call as an isolated step, it executes complete Python or R code cells in persistent kernels while a separate control plane handles orchestration, permissions, monitoring, and audited services. An append-only Action Ledger, per-cell execution records, versioned artifacts, environment manifests, and workspace checkpoints make long-running studies inspectable, branchable, and partially recoverable. The system supports model interfaces from OpenAI, Anthropic, and Gemini, and includes 604 scientific Skills. Evaluation covered 36 scenarios across retrosynthesis, molecular dynamics, protein binder design, protein mutation, catalyst screening, and mineral spectroscopy. OpenAI4S achieved a macro score of 7.83, compared with 5.79 for Kimi-K3, 6.36 for GLM-5.2, and 6.12 for Opus-4.8 running through the general-purpose Claude Code harness. Its strongest gains appeared in computation-heavy, long-horizon protein workflows, where preserving intermediate evidence and final reports mattered as much as producing plausible code. However, reproducibility remains incomplete: in-memory state is lost after kernel restarts, dependency and lineage capture is limited, and no evaluated system guarantees full rerunnability. The result is chiefly an advance in scientific workflow auditability, not proof that an AI-generated scientific conclusion is correct.

Original abstract

AI co-scientists could accelerate computational research, but over a long-running study the workflow also has to stay inspectable, resumable and reproducible, which requires persistent computational state and provenance. Here we present OpenAI4S, an open-source scientific research agent built around the principle of \emph{Code as Action, Science as Sessions}. OpenAI4S combines a persistent computing runtime with research-session management: orchestration is handled through structured tool calls, while scientific actions are represented as complete code cells executed in persistent Python and R kernels. An append-only Action Ledger, per-cell execution records, versioned artifacts, environment records, and workspace checkpoints preserve how results were produced and support session recovery, branching, and extension. Configurable sandboxing, permission controls, and code and trajectory screening provide complementary safeguards. We evaluate OpenAI4S on 36 research scenarios spanning retrosynthesis, molecular dynamics, protein binder design, protein mutation, catalyst screening, and mineral spectroscopy, measuring scientific task accuracy, workflow completeness, and reproducibility of the resulting repositories. OpenAI4S achieves an overall score of 7.83, compared with 5.7--6.4 for a general-purpose coding harness evaluated with three frontier models, with the largest gains on long-horizon and computation-intensive workflows. These results suggest that integrating persistent execution with session-level provenance can improve the reliability of AI-assisted scientific workflows. Environment specification and full rerunnability remain weak for every evaluated system, ours included, so reproducibility is still an open problem for scientific agents. The system is available under the MIT license at \href{https://github.com/PKU-YuanGroup/OpenAI4S}{github.com/PKU-YuanGroup/OpenAI4S}.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis