NTH

Atria Dawn: The Dawn of Agentic Superintelligence

AuthorsHonglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu, Zixian Huang, Jiajie Jin, Alexander Lam, Yining Li, Jiahang Lin, Yanjiang Liu, Xinyu Lu, Haijun Lv, Junlin Shang, Qisheng Su, Guoqiang Wang, Rui Wang, Zhecan Wang, Hao Xiang, Xinchen Xie, Shuhao Xing, Xiaoyu Xing, Wanghan Xu, Xinyu Yang, Yajie Yang, Chengfeng Zhao, Haoran Zhao, Ruojun Zhou, Yunhua Zhou, Yicheng Zou, Kun Cai, Qiye Cai, Xinmeng Che, Haodong Chen, Jiabei Chen, Jiahao Chen, Jiayi Chen, Yujia Chen, Lizhi Cui, Youheng Dai, Xin Deng, Yi Dong, Shihan Dou, Chenya Gu, Xu Guo, Ding Han, Feiyang Hao, Haotan He, Jie Hou, Binze Hu, Zijian Hu, Junhao Huang, Huicheng Jiang, Jiazhen Jiang, Shufan Jiang, Jiahao Kuang, Bowen Lai, Bo Li, Jiaqiang Li, Peng Li, Qilong Li, Zhuoqun Li, Jiaxiang Liu, Shuainan Liu, Tong Liu, Yi Liu, Zhonghang Lu, Jianwen Luo, Yanyi Luo, Huijie Lv, Ningsheng Ma, Zerun Ma, Houcheng Min, Chengjun Pan, Qiyuan Peng, Xiaoxuan Peng, Jianmin Qian, Jiantao Qiu, Wanying Ren, Huayu Sha, Jifei Shan...

September 20, 2026 3 min read
Watch on YouTube
The one-line take

Atria Dawn explores how research-oriented AI agents can move beyond completing tasks to partnering with humans on scientific discovery while preserving human oversight.

Key results

744B
Foundation model size

Mixture-of-experts foundation model underlying Atria Dawn Preview

16
Evaluation breadth

Benchmarks spanning research, digital work, engineering, and cybersecurity

5
Highest benchmark scores

Benchmarks on which Atria Dawn achieved the highest reported score

86.5
CyberGym score

Highest reported result on the cybersecurity benchmark

33.2%
Tasks infeasible without AI

Completed AI-assisted tasks participants judged infeasible under fixed constraints

85.5%
Human final selection

Method or parameter decisions in which humans made the final choice

What the paper found

Atria Dawn Preview is a foundation agentic language model built on a 744B-parameter mixture-of-experts architecture and trained with a Verifiable Experience Pipeline, which links tool use, execution artifacts, and externally verified outcomes. Across 16 benchmarks covering research, browsing, workspace operations, software engineering, machine-learning engineering, and cybersecurity, it records the highest reported score on 5 benchmarks, including 96.0 on DeepSearchQA, 92.5 on BrowseComp, and 86.5 on CyberGym, competing with systems such as DeepSeek, GPT, Claude, and Qwen. The paper’s central novelty is its analysis of AI-assisted AI development: agents increasingly generate methods, run experiments, diagnose failures, and implement revisions, while humans retain responsibility for goals, evaluation criteria, and final choices. In 769 task records, AI was used in 96.5% of tasks with definitive usage reports, and 33.2% of completed AI-assisted tasks were judged infeasible without AI under the same constraints. Humans made the final choice in 85.5% of method or parameter decisions, showing that greater agent execution does not equal autonomous research authority. The authors argue that recursive self-improvement requires agents to discover valuable research directions, learn intrinsic capabilities from accumulated experience, and remain subject to meaningful oversight. This complements concerns raised in work from OpenAI and Anthropic, including deployments involving Claude Code, where long-running agents still depend on human steering and accountability.

Original abstract

As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis