Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans
AuthorsAkshay K. Jagadish, Younes Strittmatter, Nori Jacoby, Eric Schulz, Nathaniel Daw, Thomas L. Griffiths, Suyog H. Chandramouli
Resources
An AI scientist discovers theories from simulated people that surprisingly predict how real humans make decisions.
Key results
Number of closed-loop theory-discovery cycles run on Centaur.
Human experiments used to test whether simulated-data theories generalized.
Mean squared error of the Relative Contextual Salience Lexicographic theory.
Mean squared error of the Contextual Relative Advantage Normalization theory.
Canonical baseline error on the held-out human experiments.
Canonical baseline error on the held-out human experiments.
What the paper found
This study tests whether cognitive theories discovered from simulated behavior can explain real human choices. Its AUTO COG system uses LLM agents to design theory-discriminating experiments, simulate responses with Centaur, arbitrate between competing executable models, and revise the weaker theory through program synthesis. The loop ran for 25 cycles in a multi-attribute decision task involving cardinal-valued cues. Centaur, a behavioral foundation model trained on trial-level data from 160 experiments, was implemented as Llama-3.1-Centaur-70B, while the discovery agents used Gemini-3.1-pro-preview; inference ran on an NVIDIA H200 GPU. The system discovered two mechanisms absent from its seed theories: Relative Contextual Salience Lexicographic, or RCSL, which combines validity-ordered search with probabilistic stopping and compensatory integration, and Contextual Relative Advantage Normalization, or CRAN, which normalizes cue differences by the strongest difference in the current context. On 10 held-out human experiments, RCSL achieved mean squared error 0.021 and CRAN 0.034, outperforming the canonical Take-The-Best score of 0.110 and Tallying score of 0.161. Their performance nearly matched theories discovered with humans in the loop. The result suggests that a simulator need not reproduce human behavior exactly: it mainly needs to preserve the distinctions that separate candidate theories. Synthetic discovery is therefore promising for broadening theory search, while human experiments remain necessary for final validation, especially when tasks move beyond the simulator’s training distribution.
Original abstract
Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is unclear whether theories discovered on them generalize to humans or merely characterize the simulator. We ran the Automated Cognitive Scientist (\textsc{AutoCog}), a closed-loop discovery system in which LLM agents design theory-discriminating experiments, collect responses, arbitrate between competing theories, and synthesize successors, entirely on behavior simulated by Centaur, a foundation model of human behavior. In a multi-attribute decision-making setting, the theories \textsc{AutoCog} found on Centaur generalized to human data: they outperformed canonical theories on ten held-out experiments and were rivaled only by theories found by running the same loop on people. We argue that this succeeds despite the simulator's inevitable imperfections because a discovery loop that arbitrates between competing theories demands less of its simulator than estimation does. The simulator only needs to capture the regularities that distinguish the theories, and not necessarily reproduce behavior precisely. Imperfect simulators can therefore widen the search over theories, with human data then testing whether the surfaced theories generalize.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.