NTH

Environment-free Synthetic Data Generation for API-Calling Agents

AuthorsSeanie Lee, Sanjoy Chowdhury, Chao Jiang, Cheng-Yu Hsieh, Ting-Yao Hu, Alexander T Toshev, Oncel Tuzel, Raviteja Vemulapalli

July 21, 2026 2 min read
Watch on YouTube
The one-line take

This work trains API-calling agents without building real backends by having LLMs simulate entire interactive worlds and generate useful synthetic trajectories.

Key results

9K
AppWorld ESAT trajectories

Generated from seven real AppWorld apps and used for fine-tuning.

6K
Synthetic-app ESAT trajectories

Generated from 52 synthetic apps spanning 1,017 APIs.

1.7K
OfficeBench ESAT trajectories

Generated from OfficeBench’s 20 APIs across eight apps.

50.5%
Maximum AppWorld gain

Improvement from ESAT fine-tuning on an AppWorld test split.

60.5%
Maximum OfficeBench gain

Improvement from ESAT fine-tuning on an OfficeBench split.

93.7%
Simulator response validity

Responses judged valid by OpenAI GPT-5.1 for schema, arguments, and state consistency.

What the paper found

Researchers at Apple introduce ESAT, an Environment-Free Synthetic Agentic Trajectory pipeline that generates API-calling training data from API specifications alone, without executable services, databases, or human-written tasks. A GLM-4.7-FP8 task generator creates coverage-balanced, intent-level tasks; a GLM-5.1-FP8 teacher agent solves them step by step while a stateful simulator produces schema-compliant responses conditioned on the task and per-app interaction history; then Gemini-3.1-Pro from Google DeepMind filters trajectories for correctness and completeness. The authors evaluate supervised fine-tuning on AppWorld and OfficeBench, including write operations and long-horizon multi-app workflows. ESAT produces 9K AppWorld trajectories from seven real AppWorld apps, 6K trajectories from 52 synthetic apps spanning 1,017 APIs, and 1.7K OfficeBench trajectories. Fine-tuning yields gains up to 50.5% on AppWorld and 60.5% on OfficeBench, with synthetic-app data often outperforming trajectories collected in the real AppWorld environment and enabling smaller Qwen3 and Qwen3.5 models to rival larger systems such as GPT-4o and Nemotron-3-120B. Quality analysis using OpenAI’s GPT-5.1 finds 93.7% of simulated API responses valid, while Gemini-3.1-Pro achieves 95.2% precision against AppWorld’s executable verifiers. The central result is that an LLM can serve as a persistent digital world model, making scalable, stateful API-agent supervision a specification problem rather than an infrastructure project.

Original abstract

Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic, pre-populated backend databases, creating a major bottleneck for scalability. To overcome this, we propose an environment-free synthetic data generation approach that leverages LLMs as on-the-fly digital world models. Given only API specifications, our method generates trajectories mimicking interactions between an agent and a stateful environment. Specifically, an LLM first generates diverse tasks solvable with the provided APIs. A teacher agent then iteratively solves each task while an LLM simulator generates coherent synthetic API responses conditioned on the task context and simulation history. Finally, an LLM judge filters the trajectories to ensure the quality of the resulting dataset. We evaluate our approach on the challenging AppWorld and OfficeBench benchmarks, which include both information-retrieval and state-changing tasks. Fine-tuning models on our synthetic data yields significant performance gains, demonstrating that effective supervision for API-calling agents can be generated without any executable environment. Our results establish LLM-based API simulation as a practical, scalable solution for training agents across diverse API ecosystems.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis