NTH

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

AuthorsXinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu

AffiliationsGeorgia Institute of Technology · Alibaba Token Foundry, Alibaba Group

October 3, 2026 2 min read
Watch on YouTube
The one-line take

VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.

Key results

3300
Admitted environments

Number of solver-grounded agentic environments generated and admitted.

0.204
Mean agentic score before training

Baseline mean across five optimization families.

0.815
Mean agentic score after training

Qwen3.6-35B-A3B score after GRPO training.

2.84
BFCL V4 improvement

Point gain across ten interaction-focused cells.

3.4
E-Commerce Bench balance multiplier

Trained model’s mean ending balance relative to the base over 365 days.

What the paper found

VHD-Play introduces mechanism-first construction for agentic reinforcement learning: instead of generating an environment and inventing its evaluator afterward, it samples and solves a formal mechanism first, then uses the solution to define both executable state dynamics and a verifier-backed outcome signal. A frozen Qwen3.6-35B-A3B setter grounds optimization problems such as inventory control, routing, knapsack, and scheduling in passages from 28 topical domains, while asymmetric tools hide parameters and require information gathering before consequential decisions. The pipeline admitted 3,300 environments, each costing roughly $0.01–$0.03, and trains policies with GRPO using terminal rewards normalized between a default policy and a solver-computed optimum. Training Qwen3.6-35B-A3B raises mean agentic performance across five optimization families from 0.204 to 0.815 while preserving written-out problem solving near 0.992. Improvements transfer to unseen mechanism families and external benchmarks: BFCL V4 rises by 2.84 points, TravelBench increases from 0.700 to 0.794, and on E-Commerce Bench the trained model reaches 3.4× the base ending balance over 365 days, outperforming Qwen3.7-Max. Comparisons between written-out, parameter-revealed, and parameter-hidden settings show that most of the learnable deficit comes from stateful interaction—planning across horizons, probing efficiently, and managing persistent commitments—rather than from solving the underlying mathematics.

Original abstract

Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc. VHD-Play reverses this dependency by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. The executable dynamics and trajectory-scoring reference are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic. Gains also appear on held-out instances from all three training families and eight unseen mechanism families, then extend beyond the generated substrate to external benchmarks for general function calling, travel planning, and 365-day e-commerce. On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max. We compare written-out problems with stateful versions that reveal or hide their parameters. The comparison shows that most of the learnable gap lies in stateful interaction rather than underlying problem solving. A frozen 35B setter realizes larger environments, and scale-matched training retains gains as mechanism size and horizon grow, indicating the potential for an evolving training substrate.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →