NTH

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

AuthorsDeyao Zhu, Xin Zhou, Shengling Qin, Xuekai Zhu, Hangliang Ding, Shu Zhong, Zixin Wen, Zhonglin Xie, Chenhui Gou, Linxuan Ren, Yueyang Wang, Junfeng Zhong, Rui Liu, Tian Gao, Yangguang Lin, Jingyuan Zhang, Maojia Song, Xuan Qi, Jinhong Wu, Chenyang Zhang, Yinzhu Piao, Ziru Niu, Hongbin Lin, Lingxiang Meng, Peng Tang, Chengyao Tang, Shanyu Wu, Huanyu Zheng, Yu Liu, Liya Zhu, He Wang, Ming Ding, Ziyu Wan, Hao Liu, Sibo Wang, Haotian Zhu, Xintian Zhang, Nan Chai, Yipeng Liu, Panhao Lai, Sihang Yuan, Zixin Su, Ge Zhang, Wangchunshu Zhou, Yantao Du, Wenhao Huang, Guang Shi

July 10, 2026 2 min read
Watch on YouTube
The one-line take

EdgeBench is a large real-world benchmark that reveals a surprising scaling law for how agents improve from long-horizon interaction after deployment.

Key results

134
tasks

real-world tasks in EdgeBench

38000
interaction_hours

total agent-environment interaction hours

0.998
mean_r2_full_benchmark

log-sigmoid fit over all 134 tasks

0.993
long_horizon_r2

log-sigmoid fit over 28-hour and 72-hour horizons

0.997
forecast_r2

predicting 12-hour performance from the first 6.5 hours

51.3
top_12h_score

Claude Opus 4.8 overall 12-hour score

What the paper found

EdgeBench from ByteDance Seed introduces a benchmark for studying post-deployment learning in real-world environments, spanning 134 executable tasks across scientific problems and ML, systems and software engineering, combinatorial optimization, professional knowledge work, formal math, and games. Across roughly 38,000 hours of agent interaction, the paper reports that aggregate best-so-far performance over time follows a highly precise log-sigmoid law, S(t)=Smax/(1+(tmid/t)^β), with mean R²=0.998 over the full 134-task average and R²≥0.993 even when extending trajectories to 28 and 72 hours. The same functional form fits each task family, outperforms log-probit, log-Gompertz, Weibull, and log-linear baselines, and can forecast 12-hour performance from only the first 6.5 hours with R²≥0.997 and RMSE below 1.0. The authors also argue for a frontier-expansion mechanism on latent task graphs, where accumulated unlocked capability and remaining locked opportunity drive sigmoid growth in log time. On the model side, learning speed roughly doubles every 3 months across frontier systems, rising by about 8× from GPT-5-Codex in September 2025 to GPT-5.5 in April 2026. In the 12-hour leaderboard, Claude Opus 4.8 leads overall with 51.3, followed by GPT-5.5 at 48.4, and a stateful 12-hour run on the gravitational-wave reconstruction task reaches 67.0 after 247 scored evaluations.

Original abstract

Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis