NTH

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

AuthorsMarkus J. Buehler

July 29, 2026 3 min read
Watch on YouTube
The one-line take

The study tests whether a language model genuinely tracks materials physics by reading and manipulating its internal representations rather than judging answers alone.

Key results

39
Constitutive-law orientation

Correctly oriented 39 of 40 directional laws in the 60-law benchmark

12
Grain-size steering

Physically correct answer shifts in all 12 of 12 refinement/coarsening conditions

9
Blinded mechanism-family identification

OpenAI gpt-5.5 identified 9 of 10 mechanism families from decoded word sets

0.642
Natural-question graph AUC

Option-free hidden-state graph performance at the natural question boundary

42
Gemma transformer layers

google/gemma-4-E4B-it model architecture

What the paper found

MIT researcher Markus J. Buehler investigates whether a language model’s correct materials-science answer reflects physical reasoning or merely lexical and numerical shortcuts. Using Google’s open-weight google/gemma-4-E4B-it, a 42-layer transformer, the study combines direct unembedding, three independently fitted Jacobian lenses, hidden-state geometry, counterfactual comparisons, and causal activation interventions. In 50 held-out descriptions spanning ten mechanism families, target-free decoded word sets enabled blinded identification of 9 of 10 families, although direct and Jacobian readouts performed similarly and prompt wording remained a strong confound. A 72-prompt graph showed comparative organization at the natural question boundary, with AUC 0.642, but exact audits found that its apparent physical polarity was equally explained by numerical direction. The strongest representational result came from a neutral-anchored 60-law benchmark: matched hidden-state changes distinguished direct, inverse, and neutral constitutive relations, correctly orienting 39 of 40 directional laws, while lexical controls were near chance. Causal steering was narrower but compelling: one grain-size direction reversed appropriately between refinement and coarsening in all 12 of 12 matched conditions, yet failed when transferred to new answer vocabularies and the inverse Hall–Petch regime. Full-state patching transferred late numerical and decision features across mechanisms, but not a universal physical-relation representation. An OpenAI gpt-5.5 interpreter recognized 9 of 10 mechanism families from decoded word sets, reinforcing semantic readability without proving causal use. The paper’s central conclusion is that materials relationships are more reliably revealed by controlled state transformations than by absolute hidden-state geometry, and that causal interpretability remains context-dependent.

Original abstract

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis