TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL
AuthorsTianze Yang, Yucheng Shi, Ruitong Sun, Jingyuan Huang, Ninghao Liu, Jin Sun
Resources
TRON is a scalable online training environment that generates fresh, automatically verifiable visual reasoning tasks on demand, helping multimodal models learn harder reasoning skills more efficiently.
Key results
rule-verifiable generator–verifier environments in the TRON suite
model-free environment audit probes
overall generation success rate in the audit
Qwen3-VL-4B-Instruct base-model pass rate at difficulty level 0
Qwen3-VL-4B-Instruct base-model pass rate at difficulty level 9
mean benchmark score after TRON, up from 52.61
What the paper found
TRON, from the University of Georgia, reframes visual reasoning reinforcement learning as training on online, rule-verifiable environments instead of static VQA corpora. The paper introduces 520 generator–verifier programs grouped into five ability buckets—spatial, mathematical, diagram, pattern/logic, and counting—each with a 0–9 difficulty ladder that can stream fresh image-question rollouts on demand and return exact rewards without an LLM judge. A model-free audit over 8,320 probes reports 99.07% generation success, 502 of 520 environments graded A, and a difficulty curve on Qwen3-VL-4B-Instruct that drops from 72.8% pass rate at level 0 to 41.3% at level 9, confirming that the curriculum axis is real rather than nominal. Using DAPO-style RL on Qwen3-VL-4B-Instruct, Qwen2.5-VL-7B-Instruct, and MiMo-VL-7B-SFT, TRON improves mean accuracy on ten external multimodal reasoning benchmarks from 52.61 to 55.23, 40.85 to 43.35, and 63.37 to 66.50 respectively. The specialist analysis shows that capability transfer matters more than visual format alone: a math specialist gains +20.0 on MM-HELIX maze and a diagram specialist gains +10.0 on PuzzleVQA rect-height, but the visually aligned specialist wins only 1 of 10 benchmarks, so robust transfer comes from covering the underlying reasoning operations, not just matching the surface modality.
Original abstract
Reinforcement learning (RL) for visual reasoning needs scalable, verifiable, and controllable training signals. Existing visual RL post-training trains on static curated datasets, with fixed image-question-answer samples bounded by their collection budget. In this work, we introduce TRON (Targeted, Rule-verifiable Online eNvironments), an online environment substrate: a training rollout is generated on demand by a controllable generator-verifier program that samples a fresh latent visual state, renders an image, asks a question, and exactly verifies the answer. A single run can therefore draw an unbounded stream of fresh instances at the difficulty level required by the current curriculum. The current TRON suite contains 520 environments organized into five ability buckets (spatial, mathematical, diagram, pattern/logic, and counting); the same substrate supports both a single full model trained on all buckets and per-bucket ability-specialist models, with no additional data collection. We also introduce a substrate analysis covering generation reliability, instance and level diversity, cross-environment near-duplicates, and base-model pass rate by difficulty level. RL post-training with METHOD consistently improves performance on ten external multimodal reasoning benchmarks across Qwen3-VL-4B, Qwen2.5-VL-7B, and MiMo-VL-7B-SFT.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.