Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
AuthorsMengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
Resources
Mechanist turns AI into an autonomous scientific instrument for discovering, testing, and controlling the mechanisms behind intelligent behavior.
Key results
Mechanist grounds hypothesis generation in an interpretability-focused graph of approximately 13K papers.
The broader knowledge base spans 43M papers across 26 scientific fields.
Mechanist achieved 92.2% human-rated reliability for experiment execution across reproduced research.
A Qwen3.5-9B student trained on filtered safe data reached a 48.6% unsafe-response rate.
Mechanism-guided intervention improved held-out belief reasoning on Pythia-410M by 15.3%.
Targeted feature steering raised mean predicted alpha-helical content to 56.6% across 900 generated sequences.
What the paper found
Mechanist is an agentic scientific instrument for studying AI models rather than merely optimizing their outputs. It cycles through hypothesis generation, experiment execution, verification, and revision, grounding its proposals in an interpretability graph of 13K papers, a cross-disciplinary corpus of 43M papers across 26 fields, and 32 mechanistic methods including sparse autoencoders, Fisher-based localization, causal ablation, probing, and activation steering. Against Claude Code and AI Scientist, it achieved a human-rated experiment-execution reliability of 92.2% across 16 paper reproductions. Its discoveries include a multimodal subliminal-learning risk: a Qwen3.5-9B student trained only on GPT-4o-filtered safe laboratory content produced unsafe answers at 48.6%, versus 20.3% for the untuned baseline. Mechanistic analysis of Pythia models identified separable Personal Belief and Attributed Belief attention heads, whose causal modulation improved held-out belief reasoning by 15.3% on Pythia-410M, without additional model training. The same approach used sparse-autoencoder features to steer Evo2-7B DNA generation, raising predicted alpha-helical content from 43.8% to 56.6% across 900 sequences while retaining structural evaluation. The framework also draws on GPT-5.4 for iteration, DeepSeek-V3.2-Thinking for knowledge-graph extraction, and tests behavioral generality across GPT, Claude, Gemini, and Qwen models, positioning mechanistic interpretability as an iterative method for discovering, auditing, explaining, and controlling intelligence.
Original abstract
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.