Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
AuthorsYijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
Resources
Business Arena tests whether LLM agents can actually run a profitable cross-border business and finds that even leading models still trail human-designed strategies.
Key results
Frontier models from families including Gemini, GPT, Claude, DeepSeek, Qwen, Kimi, and MiniMax.
Gemini 3.1 Pro's average final net worth in dollars.
MiniMax M2.5's average final net worth in dollars.
Fold difference between the highest and lowest model mean final net worth.
Mean final net worth achieved by the strongest expert-designed strategy.
Share of evaluation runs that ended below the $80,000 starting capital.
What the paper found
Business Arena introduces a controlled benchmark for long-horizon business intelligence, placing an LLM agent in a cross-border B2B marketplace grounded in Alibaba.com sourcing data, Google Trends, tariff records, seasonal demand, supplier delays, competitor repricing, customer service, finance, and compliance. Agents operate autonomously for 30 simulated days through more than 60 tools, managing the full loop from product selection and sourcing to pricing, advertising, fulfillment, recovery, and capital allocation. The benchmark evaluates 15 frontier models, including Gemini, GPT, Claude, DeepSeek, Qwen, Kimi, and MiniMax, using final net worth alongside skill-level diagnostics and action-level attribution. Results show a wide performance gap: Gemini 3.1 Pro averages $188,488, while MiniMax M2.5 averages $20,856, a 9.0-fold difference; 51% of runs lose money relative to the $80,000 starting capital. The strongest expert-designed strategy reaches $436,195, demonstrating substantial headroom beyond current models. Behavioral analysis reveals distinct operating styles: Gemini 3.1 Pro protects a 52.0% average order margin, while GPT-5.6 Sol achieves more than 93% sell-through by operating like a volume wholesaler. Mechanism ablations show that evidence-guided sourcing, full-cost pricing, tariff-aware routing, event verification, and margin-safe customer service outperform blind or shortcut policies, supporting the claim that scores measure business reasoning rather than simulator exploitation. Stateful save–fork–load evaluation further enables controlled counterfactuals and test-time trace search, making the benchmark useful for diagnosis, credit assignment, and future agent training.
Original abstract
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce \textbf{Business Arena}, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.