NTH

EvoGraph: Hybrid Directed Graph Evolution toward Software 3.0

AuthorsIgor Costa, Christopher Baran

June 23, 2026 2 min read
Watch on YouTube
The one-line take

EvoGraph is a new system that lets software evolve its own code, docs, and pipelines using a graph-based mutation-and-selection loop powered by small language models.

Key results

83%
security vulnerabilities fixed

Known vulnerabilities fixed on the evaluation benchmarks

93%
COBOL to Java functional equivalence

Test-verified translation result

90%
computational cost reduction

SLM-based modernization cost versus large language models

58%
manual review edit reduction

Transmute operator on 1.2K COBOL programs

12
GPU hours for 10K LOC

EvoGraph resource use in the cost comparison

What the paper found

EvoGraph, from AutoHand AI, proposes Hybrid Directed Graph Evolution as a closed-loop modernization system that treats code, build pipelines, documentation, tickets, schemas, logs, and runtime telemetry as one typed directed graph, then mutates that graph with small-language-model-driven operators and a multi-objective safety gate. The framework combines weight merging, AST patching, documentation sync, build weaving, and cross-language transmutation, while ranking candidates with a contextual bandit and Pareto-plus-novelty selection. On seven legacy-style benchmarks, EvoGraph fixed 83% of known security vulnerabilities, translated COBOL to Java with 93% functional equivalence, and kept documentation freshness within two minutes; across multi-language modernization it reached 82–96% semantic equivalence on COBOL, .NET, Lisp, CGI, ColdFusion, legacy Python, and C, while reducing computational cost by 90% versus large language models. The system also cut p95 latency by 40% and feature lead time by 7× relative to strong baselines, and on 1.2K COBOL programs its transmutation step reduced manual review edits by 58%. In resource terms, processing 10K lines of code required 12 GPU hours and 45 kWh, compared with 120 GPU hours and 450 kWh for a GPT-4-based approach, highlighting the paper’s central claim that specialized small language models make production-safe Software 3.0 style evolution practical rather than purely aspirational.

Original abstract

We introduce **EvoGraph**, a framework that enables software systems to evolve their own source code, build pipelines, documentation, and tickets. EvoGraph represents every artefact in a typed directed graph, applies learned mutation operators driven by specialized small language models (SLMs), and selects survivors with a multi-objective fitness. On three benchmarks, EvoGraph fixes 83% of known security vulnerabilities, translates COBOL to Java with 93% functional equivalence (test verified), and maintains documentation freshness within two minutes. Experiments show a 40% latency reduction and a sevenfold drop in feature lead time compared with strong baselines. We extend our approach to **evoGraph**, leveraging language-specific SLMs for modernizing .NET, Lisp, CGI, ColdFusion, legacy Python, and C codebases, achieving 82-96% semantic equivalence across languages while reducing computational costs by 90% compared to large language models. EvoGraph's design responds to empirical failure modes in legacy modernization, such as implicit contracts, performance preservation, and integration evolution. Our results suggest a practical path toward Software 3.0, where systems adapt continuously yet remain under measurable control.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis