Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
AuthorsBytedance Seed
Resources
Seed2.0 is a large foundation model series aimed at improving real-world reasoning, instruction following, and visual understanding, presented through a model card focused on practical deployment and evaluation.
Key results
Seed2.0 Pro score on AIME 2025
Seed2.0 Pro Codeforces Elo rating
Seed2.0 Pro score on Putnam-200
Seed2.0 Pro score across five ICPC contests
Seed2.0 Pro overall score on the 912-case Chinese complex instruction benchmark
Seed2.0 Pro score on long-video understanding
What the paper found
Seed2.0 Model Card, released by ByteDance Seed, presents a production-oriented model family that spans Pro, Lite, and Mini variants and is explicitly optimized for real-world complexity rather than isolated benchmark wins. The paper argues that the key bottlenecks in deployed AI are multimodal understanding, low-latency inference, reliable complex instruction following, and long-horizon agentic execution, then evaluates Seed2.0 across frontier math, coding, vision, video, search, and workflow benchmarks against GPT-5.2, Claude-Opus-4.5, Claude-Sonnet-4.5, and Gemini-3 Pro/Flash. Seed2.0 Pro reaches 98.3 on AIME 2025, 3020 Codeforces Elo, 35.5 on Putnam-200 Pass@8, and 73.02% Pass@8 on five ICPC contests, while also posting 75.26 on a 912-case Chinese complex instruction benchmark, improving over Seed1.8 by 2.37 points. In multimodal evaluation it sets state-of-the-art scores such as 88.8 on MathVision, 90.5 on MathKangaroo, 92.3 on DA-2K, 72.4 on DUDE, and 89.5 on VideoMME, and it strengthens agentic search and research performance with 77.4 on DeepSearchQA and 50.8 on ResearchRubrics. The model card also emphasizes cost efficiency: Seed2.0 Pro, Lite, and Mini are priced at $0.47, $0.09, and $0.03 per 1M input tokens, with output pricing as low as $0.31 per 1M tokens for Mini, positioning the series as a lower-cost alternative for large-scale enterprise deployment. Case studies show Seed2.0 solving cryptanalysis, repository construction, GUI automation, and scientific reasoning tasks, including 231061.93 mm³ volume and 23306.19 mm² surface-area verification in FreeCAD, underscoring the paper’s central claim that the model family is moving toward agentic, tool-using intelligence for production environments.
Original abstract
We present Seed2.0, a model series that takes a meaningful step toward solving complex, real-world tasks. Our approach begins with identifying users' genuine needs and constructing a reliable, forward-looking evaluation system by selecting and abstracting benchmarks grounded in these needs and in realistic, complex scenarios. Guided by this evaluation system, Seed2.0 targets two persistent challenges, long-tail knowledge and complex instruction following, substantially improving the model's reliability on intricate, long-horizon tasks. Beyond these, Seed2.0 delivers world-leading reasoning intelligence, visual understanding, and search capabilities that address the most common needs of a broad user base. Through extensive real-world use cases documented in this model card, we demonstrate that Seed2.0 begins to exhibit the ability to handle initial complex real-world tasks, delivering greater value to hundreds of millions of users.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.