NTH

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills

AuthorsChuan Xiao, Zhengbo Jiao, Shaobo Wang, Wei Wang, Bing Zhao, Hu Wei, Linfeng Zhang, Lin Qu

June 21, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how coding agents can learn from their own past mistakes by turning old solving traces into new training tasks, steadily improving software-engineering performance over multiple iterations.

Key results

50.40%
SWE-bench Verified

best result after 3 iterations

36.67%
SWE-bench Lite

best result after 3 iterations

22.85%
SWE-bench Pro

best result after 3 iterations

14.61%
Terminal-Bench 2.0

best result after 3 iterations

100
validation set size

held-out BeyondSWE tasks used for gradient alignment

What the paper found

Socratic-SWE from Alibaba Group and Shanghai Jiao Tong University proposes a closed-loop self-evolving framework for software engineering agents that turns historical solving traces into an Agent Skill Registry, then uses those skills to generate repository-grounded repair tasks, filter them with execution-based validation, and rank them by solver-gradient alignment rather than raw difficulty. The system is built on Qwen3.5-9B for both Generator and Solver, with Qwen3.6-27B used for skill distillation, and it trains for 3 iterations on 12,000 validated instances per iteration, using 36,000 total training instances and a held-out validation set of 100 BeyondSWE tasks. Under the same compute budget, Socratic-SWE reaches 50.40% on SWE-bench Verified, 36.67% on SWE-bench Lite, 22.85% on SWE-bench Pro, and 14.61% on Terminal-Bench 2.0, outperforming self-evolving baselines such as SSR, Socratic-Zero, Absolute-Zero, SPIRAL, and R-Zero. The paper’s main ablation shows the Skill Registry is the largest driver of gains, while cosine-based gradient alignment beats an inner-product scorer, and the system scales to 52.00% on SWE-bench Verified after 5 iterations before saturating.

Original abstract

LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by the availability of high-quality SWE tasks. Existing synthetic data methods typically create tasks through fixed mutation or bug-injection procedures, making the resulting distributions largely independent of the agent's own weaknesses and training progress. We introduce Socratic-SWE, a closed-loop self-evolution framework that reuses the agent's historical solving traces as a source of training signal. Rather than treating traces only as evidence for reward computation, Socratic-SWE distills them into structured agent skills that summarize recurring failures and effective repair patterns. These skills then guide the generation of targeted repair tasks in real repositories. Candidate tasks are checked through execution-based validation and scored with a solver-gradient alignment reward, so that the retained tasks are both verifiable and useful for improving the Solver. The updated Solver produces new traces, enabling the task curriculum to adapt over successive rounds. Across SWE-bench Verified, SWE-bench Lite, SWE-bench Pro, and Terminal-Bench 2.0, Socratic-SWE consistently improves over self-evolving baselines under the same compute budget, reaching 50.40% on SWE-bench Verified after three iterations. These results suggest that solving traces can serve as a scalable substrate for self-evolving SWE agents.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis