CogEvol: Towards Efficient and Reliable Learning Environment Generation
AuthorsShangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Haoxuan Li, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang
CogEvol is a compact, production-trained model that turns course briefs into reliable slides and interactive educational web pages much faster and more cheaply than larger coding models.
Key results
Requests used to report production generation latency.
Median seconds to generate one slide.
Median seconds to generate one interactive HTML page.
Execution-verified conversations used for supervised fine-tuning.
Slide-std quality score on a 0–100 scale.
HTML-500 score on a 0–100 scale.
What the paper found
CogEvol introduces Learning Environment Generation, a post-training task in which one model converts a course brief into either renderer-valid JSON slides or self-contained interactive HTML, without multi-turn agent scaffolding. Across 220k production requests, it generates a slide in a median of 17 seconds and an interactive page in 59 seconds. Its pipeline combines 53687 execution-verified supervised examples with GRPO reinforcement learning: a hybrid rule-plus-VLM reward evaluates slide geometry and rendered fidelity, while a Playwright-driven Chromium probe measures actual interaction and applies hard-fail gates. CogEvol-27B, built on Qwen3.8-27B, reaches 83.7 on slide-std and 63.7 on the 500-case HTML-500 benchmark, with zero interactive hard failures. This contrasts with systems such as Claude Opus 4.8, GPT-5.4, Gemini 3.6 Flash, DeepSeek-V4-Pro, GLM-5.3, and Qwen3.8-Max, which often trade slide quality for broken or incomplete interactivity. The paper’s central finding is that interactivity must be measured rather than inferred from screenshots: an earlier reward-hacked model scored 18.8 on games despite visually convincing pages, while the hardened reward raised game performance to 57.6. Scaffold editing, which retrieves and patches reusable templates instead of regenerating full HTML, cuts interactive-page generation cost by about 76 percent, and the open CogEvol-4B release extends the approach to lower-cost and on-device deployment.
Original abstract
We present CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass. Across 220k production requests, CogEvol completes a slide in a median of 17 seconds and an interactive page in 59, replacing minutes-long multi-turn agent scaffolding. Reliability is enforced rather than hoped for: a production-grounded data pipeline turns real failures into 53,687 verified SFT samples, and a hybrid rule-plus-VLM reward drives GRPO-based RL, hardened after we caught and fixed a reward-hacking episode that produced visually convincing but unplayable games. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9x fewer parameters than flagship coding models, and, in collaboration with the OpenMAIC team, serves their live production traffic. CogEvol-4B is released openly under the Apache 2.0 license at https://github.com/CogEvol/CogEvol-4B; external flagships are measured on the same suites under the identical harness. Scaffold editing cuts interactive-page generation cost by a further ~76%, and the full stack runs on domestic Ascend accelerators at application-level parity with A800 GPUs, lowering the unit cost of AI-native education at scale.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.