NTH

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

AuthorsChaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang

August 27, 2026 2 min read
Watch on YouTube
The one-line take

AgentSysBench shows that serving AI agents requires systems designed for long-running, stateful, tool-using workloads rather than just faster LLM inference.

Key results

10
AgentSysBench applications

Representative agentic applications spanning retrieval, coding, browser, GUI, and research workloads.

28 GB
Sandbox working-set peak

Maximum measured per-session active sandbox memory.

32×
Task latency divergence

Maximum latency difference between tasks sharing a component.

40%
Task-disaggregated latency reduction

Best-case reduction from separating heterogeneous logical tasks into dedicated services.

4.5×
Co-location speedup

Best-case improvement from communication-aware placement of embedding and vector-database services.

35.2%
Redundant search calls removed

Reduction achieved by query-result caching with a 10-minute TTL.

What the paper found

This paper introduces AgentSysBench, a systems benchmark covering 10 agentic applications, including OpenAI Codex, Anthropic Claude Code, DeepResearch, Mini-SWE, WebAgent, and GUIAgent, with unified tracing across LLM calls, tools, sandboxes, retrieval, state, and deployment. Across 4,641 controlled requests and production traces, it shows that agent workloads are long-running, stateful, and unlike conventional single-turn LLM inference: non-LLM components dominate latency in half of the applications, sandbox working-set memory can reach 28 GB per session, and heterogeneous tasks can differ by up to 32× in latency across GPU-bound models, CPU-bound sandboxes, memory-bound vector databases, and network-bound search. The study also finds production sessions frequently remain idle but resumable, while control-plane overhead from tool schemas, observations, compaction, and auxiliary calls consumes context and computation. Cross-session redundancy creates substantial caching potential: recurring search queries and fetched URLs account for most external calls in the measured traces. Characterization-guided serving interventions demonstrate that task-disaggregated serving cuts latency by up to 40%, communication-aware co-location delivers up to 4.5× speedup, and tool-result caching removes 35.2% of redundant search calls. The central implication is that serving systems must jointly schedule models, tools, communication, and persistent state rather than optimize GPU inference alone.

Original abstract

Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state. However, the system behavior of these workloads---where latency, cost, and bottlenecks arise---remains poorly characterized, leaving serving systems to rely on assumptions built for conventional inference. We present AgentSysBench, a benchmark suite and measurement toolkit with ten representative agentic applications and unified systems-level instrumentation. Across controlled deployments and production traces, we identify six properties that distinguish agentic workloads from conventional LLM serving: (1) execution is heavyweight and stateful, with non-LLM components dominating latency in 5 of 10 applications and sandbox working-set memory peaking at 28 GB per session; (2) applications compose components with heterogeneous resource affinity---GPU-bound inference, memory-bound retrieval, CPU-bound sandboxes---whose task latencies diverge by up to 32x; (3) bottlenecks shift across requests, models, and deployments; (4) production sessions hold state idle for minutes to hours between active steps; (5) a control-plane tax---auxiliary LLM calls and context overhead from tool schemas and observations---crowds out productive compute and context; and (6) production traces from three applications reveal heavy cross-request redundancy in search queries and web fetches, exposing a large caching opportunity. Four design explorations demonstrate that these findings are actionable: task-aware serving reduces latency by 29--40%, communication-aware placement by up to 4.5x, state offloading reduces memory usage by 4.6x, and tool-result caching removes 35.2% of redundant search calls and saves 19.3% of aggregate search latency.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis