NTH

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

AuthorsIgor Itkin

August 18, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how to replace expensive LLM agents with cheap learned surrogates so large artificial societies can be simulated on an ordinary laptop.

Key results

-0.569
EconAgent Phillips correlation

Out-of-sample Phillips correlation reproduced by the cloned surrogate.

0.50
Reasoning ablation main effect

Correlation-point shift attributed to reasoning rather than wording.

0.018
Jensen curvature floor

Private-feed error floor for the saturating work response.

0.34
Laptop epidemic closure runtime

Seconds required by the degree-aware closure.

22
Epidemic closure RMSE

Mortality-curve RMSE achieved by the degree-aware closure.

What the paper found

The paper presents a low-cost alternative to running thousands of expensive LLM agents: query a teacher model such as DeepSeek a few hundred to few thousand times, fit each agent with a two-to-twelve-parameter behavioral surrogate, and simulate the resulting society on a laptop. Its central contribution is an interaction-order-by-memory taxonomy that predicts when coarse-graining works. Global feeds produce mean-field behavior, with surrogate error decreasing as N^-1/2; community feeds require block or graphon closures and leave an O(1) floor; local, long-memory systems require pair approximations and memory kernels. Across eight named simulations, including EconAgent, OASIS, AgentSociety, LLMTraveler, Generative Agents, and epidemic models, the predicted error trends held. In EconAgent, the surrogate reproduced the behavioral Phillips correlation at -0.569, while showing that the reported Okun relationship is largely an accounting identity. A 2×2 ablation found that reasoning, rather than prompt wording, shifts the Phillips effect by 0.50 correlation points. On a genuine DeepSeek response function, curvature created a private-feed error floor of 0.018, demonstrating when Jensen bias defeats averaging. The method also transfers beyond LLM societies: a degree-aware epidemic closure achieved RMSE 22 in 0.34 seconds, versus RMSE 48 and roughly 400 seconds for a differentiable GPU model. Cross-model tests spanning DeepSeek, OpenAI’s GPT-4o, Anthropic, Google, Meta, and Llama suggest that perception structure is broadly shared, while response curvature and reasoning remain model-specific.

Original abstract

Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis