NTH

CogEvol: Towards Efficient and Reliable Learning Environment Generation

AuthorsShangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Haoxuan Li, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang

September 5, 2026 2 min read
Watch on YouTube
The one-line take

CogEvol is a compact, production-trained model that turns course briefs into reliable slides and interactive educational web pages much faster and more cheaply than larger coding models.

Key results

220k
Production requests

Requests used to report production generation latency.

17
Slide median latency

Median seconds to generate one slide.

59
Interactive-page median latency

Median seconds to generate one interactive HTML page.

53687
Verified SFT samples

Execution-verified conversations used for supervised fine-tuning.

83.7
CogEvol-27B slide score

Slide-std quality score on a 0–100 scale.

63.7
CogEvol-27B HTML score

HTML-500 score on a 0–100 scale.

What the paper found

CogEvol introduces Learning Environment Generation, a post-training task in which one model converts a course brief into either renderer-valid JSON slides or self-contained interactive HTML, without multi-turn agent scaffolding. Across 220k production requests, it generates a slide in a median of 17 seconds and an interactive page in 59 seconds. Its pipeline combines 53687 execution-verified supervised examples with GRPO reinforcement learning: a hybrid rule-plus-VLM reward evaluates slide geometry and rendered fidelity, while a Playwright-driven Chromium probe measures actual interaction and applies hard-fail gates. CogEvol-27B, built on Qwen3.8-27B, reaches 83.7 on slide-std and 63.7 on the 500-case HTML-500 benchmark, with zero interactive hard failures. This contrasts with systems such as Claude Opus 4.8, GPT-5.4, Gemini 3.6 Flash, DeepSeek-V4-Pro, GLM-5.3, and Qwen3.8-Max, which often trade slide quality for broken or incomplete interactivity. The paper’s central finding is that interactivity must be measured rather than inferred from screenshots: an earlier reward-hacked model scored 18.8 on games despite visually convincing pages, while the hardened reward raised game performance to 57.6. Scaffold editing, which retrieves and patches reusable templates instead of regenerating full HTML, cuts interactive-page generation cost by about 76 percent, and the open CogEvol-4B release extends the approach to lower-cost and on-device deployment.

Original abstract

We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real failures into 53,687 verified SFT samples, and a hybrid rule-plus-VLM reward drives GRPO-based RL, hardened after we caught and fixed a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol-4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol-4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive-page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application-level parity with A800 GPUs, lowering the unit cost of AI-native education at scale.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis