EvoGraph: Hybrid Directed Graph Evolution toward Software 3.0
AuthorsIgor Costa, Christopher Baran
Resources
EvoGraph is a new system that lets software evolve its own code, docs, and pipelines using a graph-based mutation-and-selection loop powered by small language models.
Key results
Known vulnerabilities fixed on the evaluation benchmarks
Test-verified translation result
SLM-based modernization cost versus large language models
Transmute operator on 1.2K COBOL programs
EvoGraph resource use in the cost comparison
What the paper found
EvoGraph, from AutoHand AI, proposes Hybrid Directed Graph Evolution as a closed-loop modernization system that treats code, build pipelines, documentation, tickets, schemas, logs, and runtime telemetry as one typed directed graph, then mutates that graph with small-language-model-driven operators and a multi-objective safety gate. The framework combines weight merging, AST patching, documentation sync, build weaving, and cross-language transmutation, while ranking candidates with a contextual bandit and Pareto-plus-novelty selection. On seven legacy-style benchmarks, EvoGraph fixed 83% of known security vulnerabilities, translated COBOL to Java with 93% functional equivalence, and kept documentation freshness within two minutes; across multi-language modernization it reached 82–96% semantic equivalence on COBOL, .NET, Lisp, CGI, ColdFusion, legacy Python, and C, while reducing computational cost by 90% versus large language models. The system also cut p95 latency by 40% and feature lead time by 7× relative to strong baselines, and on 1.2K COBOL programs its transmutation step reduced manual review edits by 58%. In resource terms, processing 10K lines of code required 12 GPU hours and 45 kWh, compared with 120 GPU hours and 450 kWh for a GPT-4-based approach, highlighting the paper’s central claim that specialized small language models make production-safe Software 3.0 style evolution practical rather than purely aspirational.
Original abstract
We introduce **EvoGraph**, a framework that enables software systems to evolve their own source code, build pipelines, documentation, and tickets. EvoGraph represents every artefact in a typed directed graph, applies learned mutation operators driven by specialized small language models (SLMs), and selects survivors with a multi-objective fitness. On three benchmarks, EvoGraph fixes 83% of known security vulnerabilities, translates COBOL to Java with 93% functional equivalence (test verified), and maintains documentation freshness within two minutes. Experiments show a 40% latency reduction and a sevenfold drop in feature lead time compared with strong baselines. We extend our approach to **evoGraph**, leveraging language-specific SLMs for modernizing .NET, Lisp, CGI, ColdFusion, legacy Python, and C codebases, achieving 82-96% semantic equivalence across languages while reducing computational costs by 90% compared to large language models. EvoGraph's design responds to empirical failure modes in legacy modernization, such as implicit contracts, performance preservation, and integration evolution. Our results suggest a practical path toward Software 3.0, where systems adapt continuously yet remain under measurable control.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.