Robotic Policy Adaptation via Weight-Space Meta-Learning
AuthorsChristian Bianchi, Siamak Yousefi, Alessio Sampieri, Andrea Roberti, Luca Rigazio, Fabio Galasso, Luca Franco
Resources
WIZARD teaches a robot policy to adapt to new tasks in one shot by predicting the right weight updates from a demo video and instruction, avoiding task-specific fine-tuning.
Key results
WIZARD zero-shot average success on held-out LIBERO-Spatial
Comparison baseline average success on LIBERO-Spatial
WIZARD average success on the Franka Emika Panda tasks
π0.5 with real-domain adaptation on the same Franka tasks
Steps for WIZARD-initialized fine-tuning to reach 96% expert success
What the paper found
Robotic Policy Adaptation via Weight-Space Meta-Learning introduces WIZARD, a weight-space meta-learning framework for vision-language-action robotics that generates task-specific LoRA adapters from a language instruction and a short demonstration video in a single forward pass, without action labels, test-time optimization, or privileged goal images. Built around the frozen π0.5 policy, WIZARD meta-trains on expert LoRA updates paired with multimodal task embeddings and adds two robotics-specific design choices: explicit per-layer scale prediction and cosine alignment in weight space. On the LIBERO benchmark, the method lifts zero-shot success on the held-out LIBERO-Spatial suite from 0.19 for the MT-VLA baseline to 0.40, while also improving LIBERO-Goal from 0.14 to 0.22 and LIBERO-Object from 0.01 to 0.03; the paper reports up to about 2× gains on unseen dataset collections and up to about 14× on unseen tasks. On a real Franka Emika Panda with three Intel RealSense cameras, WIZARD improves average success from 0.22 to 0.41 across five manipulation tasks after the same real-domain grounding, and its generated adapters can also warm-start fine-tuning to reach 96% expert success in 70 steps instead of 90. The ablation study shows that pure MSE reconstruction fails completely, scale-aware supervision raises average success to 0.27, and adding cosine alignment further increases it to 0.33, underscoring that stable parameter scale is essential for robotic weight generation.
Original abstract
Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from large corpora of demonstrations and action labels. However, adapting these models to new tasks still typically requires task-specific demonstrations, action annotations, and additional fine-tuning, making deployment costly and difficult to scale. We propose WIZARD, a weight-space meta-learning framework that sidesteps task-specific fine-tuning by generating task-specific LoRA parameters for a frozen VLA policy. Given only a language instruction and a short demonstration video, WIZARD predicts the corresponding adaptation weights in a single forward pass, without target-task action labels or test-time optimization. During meta-training, WIZARD learns to map task evidence directly to expert LoRA updates, capturing relationships between tasks in weight space. Experiments on LIBERO show that WIZARD improves performance by up to ~2x on unseen dataset collections and up to ~14x on unseen tasks. On a Franka Emika Panda, WIZARD consistently improves over a real-domain adapted baseline, showing that generated adapters provide task-level specialization beyond simulation.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.