RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
AuthorsHanan Gani, Tejal Kulkarni, Madhoolika Chodavarapu, Nicklas Hansen, Manmohan Chandraker
RoboTALES teaches robot policies by using language-guided planning and vision-language feedback to imagine task-aligned futures, improving long-horizon manipulation.
Key results
Average success rate on the RoboCasa benchmark for RoboTALES.
Average success rate on the LIBERO10 benchmark for RoboTALES.
Strongest prior baseline on LIBERO10 reported in the paper.
Reward signal calibration for predicting task success over 360 rollouts.
Jointly trained model's Fréchet distance on UCF101 future prediction evaluation.
Jointly trained model's MMD-RBF on UCF101 future prediction evaluation.
What the paper found
RoboTALES, from the University of California, San Diego, reframes robot control as task-aligned imagination by coupling a pretrained Stable Video Diffusion backbone with hierarchical reasoning and policy learning. A Gemini 2.5 Pro LLM planner decomposes each instruction into 2 to 5 ordered subgoals, the resulting plan conditions the video generator, and a frozen VLM critic provides DDPO-style reward feedback that steers the denoising dynamics toward semantically faithful futures rather than merely realistic ones. Those learned decoder features then condition a 1D action diffusion UNet, with action gradients flowing back into the video generator in a single-stage joint objective. On RoboCasa, the method reaches 64% average success across 24 manipulation tasks, improving over VideoPolicy and the strongest prior baseline, and on LIBERO10 it achieves 97% average success across 10 tasks, versus 94% for VideoPolicy. The paper reports especially strong gains on structured long-horizon behaviors such as turning, pressing, and pick-and-place, while ablations show that planner injection alone is insufficient without joint training and critic feedback. Additional analysis finds the critic reward is calibrated to success with AUROC 0.72, and joint training also improves general video generation fidelity on UCF101, reducing Fréchet distance from 0.6715 to 0.2437 and MMD-RBF from 0.2419 to 0.0631.
Original abstract
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks. Our code and models are publicly available at https://github.com/hananshafi/RoboTALES.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.