NTH

Robotic Policy Adaptation via Weight-Space Meta-Learning

AuthorsChristian Bianchi, Siamak Yousefi, Alessio Sampieri, Andrea Roberti, Luca Rigazio, Fabio Galasso, Luca Franco

June 11, 2026 2 min read
Watch on YouTube
The one-line take

WIZARD teaches a robot policy to adapt to new tasks in one shot by predicting the right weight updates from a demo video and instruction, avoiding task-specific fine-tuning.

Key results

0.40
LIBERO-Spatial avg success

WIZARD zero-shot average success on held-out LIBERO-Spatial

0.19
MT-VLA baseline on LIBERO-Spatial

Comparison baseline average success on LIBERO-Spatial

0.41
Real-world avg success

WIZARD average success on the Franka Emika Panda tasks

0.22
Real-world baseline avg success

π0.5 with real-domain adaptation on the same Franka tasks

70
Warm-start steps to expert

Steps for WIZARD-initialized fine-tuning to reach 96% expert success

What the paper found

Robotic Policy Adaptation via Weight-Space Meta-Learning introduces WIZARD, a weight-space meta-learning framework for vision-language-action robotics that generates task-specific LoRA adapters from a language instruction and a short demonstration video in a single forward pass, without action labels, test-time optimization, or privileged goal images. Built around the frozen π0.5 policy, WIZARD meta-trains on expert LoRA updates paired with multimodal task embeddings and adds two robotics-specific design choices: explicit per-layer scale prediction and cosine alignment in weight space. On the LIBERO benchmark, the method lifts zero-shot success on the held-out LIBERO-Spatial suite from 0.19 for the MT-VLA baseline to 0.40, while also improving LIBERO-Goal from 0.14 to 0.22 and LIBERO-Object from 0.01 to 0.03; the paper reports up to about 2× gains on unseen dataset collections and up to about 14× on unseen tasks. On a real Franka Emika Panda with three Intel RealSense cameras, WIZARD improves average success from 0.22 to 0.41 across five manipulation tasks after the same real-domain grounding, and its generated adapters can also warm-start fine-tuning to reach 96% expert success in 70 steps instead of 90. The ablation study shows that pure MSE reconstruction fails completely, scale-aware supervision raises average success to 0.27, and adding cosine alignment further increases it to 0.33, underscoring that stable parameter scale is essential for robotic weight generation.

Original abstract

Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from large corpora of demonstrations and action labels. However, adapting these models to new tasks still typically requires task-specific demonstrations, action annotations, and additional fine-tuning, making deployment costly and difficult to scale. We propose WIZARD, a weight-space meta-learning framework that sidesteps task-specific fine-tuning by generating task-specific LoRA parameters for a frozen VLA policy. Given only a language instruction and a short demonstration video, WIZARD predicts the corresponding adaptation weights in a single forward pass, without target-task action labels or test-time optimization. During meta-training, WIZARD learns to map task evidence directly to expert LoRA updates, capturing relationships between tasks in weight space. Experiments on LIBERO show that WIZARD improves performance by up to ~2x on unseen dataset collections and up to ~14x on unseen tasks. On a Franka Emika Panda, WIZARD consistently improves over a real-domain adapted baseline, showing that generated adapters provide task-level specialization beyond simulation.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis