CAST: Game Solvers as Turn-Level Teachers for LLM Agents
AuthorsYu Wang, Yi-Kai Zhang, Wentao Shi, Ziang Ye, Yuchun Miao, Yueqing Sun, Qi Gu, Xunliang Cai, Lan-Zhe Guo, Han-Jia Ye, Fuli Feng
CAST teaches LLM agents to make better long-horizon decisions by using a game solver’s changing state values as turn-by-turn guidance.
Key results
Average Avg@4 success rate across Sokoban, Minesweeper, and Rush Hour.
Average Avg@4 success rate on harder held-out game difficulties.
Lower end of CAST’s reported 1.7–2.0× speedup in reaching DAPO’s peak validation performance.
Overall zero-shot Avg@4 score across ALFWorld and WebShop.
Estimated parts per million of total training-step wall-clock time.
What the paper found
CAST, from researchers at the University of Science and Technology of China, Nanjing University, Wuhan University, and Meituan, addresses sparse credit assignment when training LLM agents with reinforcement learning from verifiable rewards. Instead of giving every turn the same win-or-loss signal, it queries a game solver for each LLM action, measures the change in cost-to-go, and adds this shifted solver advantage to DAPO-style GRPO training. An asinh transformation compresses harmful outliers, while batch-level RMS normalization stabilizes scales across Sokoban, Minesweeper, and Rush Hour. The authors theoretically show that, under a soft-optimal solver assumption, this scalar advantage performs logit-free on-policy distillation: it encodes the teacher’s action preference without requiring teacher logits. Using Qwen3-4B-Instruct-2507, CAST raises the trained-baseline in-domain average from 44.7 to 62.1 and the unseen-difficulty average from 18.7 to 28.4, outperforming GRPO, GSPO, DAPO, and GiGPO on every game. It reaches DAPO’s peak validation performance with a 1.7–2.0× training speedup, and transfers zero-shot to ALFWorld and WebShop with an overall score of 30.3. For context, this surpasses the prompting averages of Gemini-2.5-Flash and Claude-Sonnet-4.5 on the in-domain games. Solver queries contribute only 73 ppm of total training-step time, and a learned DQN value network retains much of the benefit of exact solver guidance.
Original abstract
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state toward success. Building on this insight, we propose CAST (Credit Assignment from Solver Teachers), which converts these value changes into solver advantages and injects them into RLVR as turn-level signals. We further show that, under a soft-optimal solver assumption, maximizing the solver advantage is equivalent to on-policy distillation from the solver, requiring only scalar values rather than teacher logits. Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines on every game under both in-domain and unseen-difficulty evaluation and achieves the highest average zero-shot performance on ALFWorld and WebShop. Our code is available at https://github.com/Wloner0809/CAST.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.