NTH

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

AuthorsHanan Gani, Tejal Kulkarni, Madhoolika Chodavarapu, Nicklas Hansen, Manmohan Chandraker

July 13, 2026 2 min read
Watch on YouTube
The one-line take

RoboTALES teaches robot policies by using language-guided planning and vision-language feedback to imagine task-aligned futures, improving long-horizon manipulation.

Key results

64%
RoboCasa average success

Average success rate on the RoboCasa benchmark for RoboTALES.

97%
LIBERO10 average success

Average success rate on the LIBERO10 benchmark for RoboTALES.

94%
VideoPolicy LIBERO10 success

Strongest prior baseline on LIBERO10 reported in the paper.

0.72
Critic AUROC

Reward signal calibration for predicting task success over 360 rollouts.

0.2437
UCF101 Fréchet distance

Jointly trained model's Fréchet distance on UCF101 future prediction evaluation.

0.0631
UCF101 MMD-RBF

Jointly trained model's MMD-RBF on UCF101 future prediction evaluation.

What the paper found

RoboTALES, from the University of California, San Diego, reframes robot control as task-aligned imagination by coupling a pretrained Stable Video Diffusion backbone with hierarchical reasoning and policy learning. A Gemini 2.5 Pro LLM planner decomposes each instruction into 2 to 5 ordered subgoals, the resulting plan conditions the video generator, and a frozen VLM critic provides DDPO-style reward feedback that steers the denoising dynamics toward semantically faithful futures rather than merely realistic ones. Those learned decoder features then condition a 1D action diffusion UNet, with action gradients flowing back into the video generator in a single-stage joint objective. On RoboCasa, the method reaches 64% average success across 24 manipulation tasks, improving over VideoPolicy and the strongest prior baseline, and on LIBERO10 it achieves 97% average success across 10 tasks, versus 94% for VideoPolicy. The paper reports especially strong gains on structured long-horizon behaviors such as turning, pressing, and pick-and-place, while ablations show that planner injection alone is insufficient without joint training and critic feedback. Additional analysis finds the critic reward is calibrated to success with AUROC 0.72, and joint training also improves general video generation fidelity on UCF101, reducing Fréchet distance from 0.6715 to 0.2437 and MMD-RBF from 0.2419 to 0.0631.

Original abstract

Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these models can be difficult to use for planning or policy extraction. To address these limitations, we propose RoboTALES, a single-stage framework that learns task-aligned simulated futures and uses them to train robot policies. Our approach introduces two key innovations: (1) a hierarchical LLM-based planner that breaks complex tasks into a sequence of subgoals to guide the model's imagination; and (2) a VLM-based critic that evaluates these ``imagined'' futures and uses reward-based feedback to keep the model's internal representations focused on the goal. By anchoring the video generator in abstract reasoning, we produce temporally consistent rollouts and more coherent actions. We evaluate RoboTALES on diverse manipulation tasks from RoboCasa and LIBERO10, and show that our method consistently outperforms existing methods, especially in long-horizon tasks. Our code and models are publicly available at https://github.com/hananshafi/RoboTALES.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis