Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
AuthorsXinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
AffiliationsGeorgia Institute of Technology · Alibaba Token Foundry, Alibaba Group
Resources
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.
Key results
Number of solver-grounded agentic environments generated and admitted.
Baseline mean across five optimization families.
Qwen3.6-35B-A3B score after GRPO training.
Point gain across ten interaction-focused cells.
Trained model’s mean ending balance relative to the base over 365 days.
What the paper found
VHD-Play introduces mechanism-first construction for agentic reinforcement learning: instead of generating an environment and inventing its evaluator afterward, it samples and solves a formal mechanism first, then uses the solution to define both executable state dynamics and a verifier-backed outcome signal. A frozen Qwen3.6-35B-A3B setter grounds optimization problems such as inventory control, routing, knapsack, and scheduling in passages from 28 topical domains, while asymmetric tools hide parameters and require information gathering before consequential decisions. The pipeline admitted 3,300 environments, each costing roughly $0.01–$0.03, and trains policies with GRPO using terminal rewards normalized between a default policy and a solver-computed optimum. Training Qwen3.6-35B-A3B raises mean agentic performance across five optimization families from 0.204 to 0.815 while preserving written-out problem solving near 0.992. Improvements transfer to unseen mechanism families and external benchmarks: BFCL V4 rises by 2.84 points, TravelBench increases from 0.700 to 0.794, and on E-Commerce Bench the trained model reaches 3.4× the base ending balance over 365 days, outperforming Qwen3.7-Max. Comparisons between written-out, parameter-revealed, and parameter-hidden settings show that most of the learnable deficit comes from stateful interaction—planning across horizons, probing efficiently, and managing persistent commitments—rather than from solving the underlying mathematics.
Original abstract
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc. VHD-Play reverses this dependency by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. The executable dynamics and trajectory-scoring reference are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic. Gains also appear on held-out instances from all three training families and eight unseen mechanism families, then extend beyond the generated substrate to external benchmarks for general function calling, travel planning, and 365-day e-commerce. On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max. We compare written-out problems with stateful versions that reveal or hide their parameters. The comparison shows that most of the learnable gap lies in stateful interaction rather than underlying problem solving. A frozen 35B setter realizes larger environments, and scale-matched training retains gains as mechanism size and horizon grow, indicating the potential for an evolving training substrate.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up
Michael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville
NGU helps RL-trained LLMs stop over-practicing easy problems and spend more effort solving the hard ones.