PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning
AuthorsYoungjoon Jeong, Jihwan Yu, Minsoo Jo, Junha Chun, Taesup Kim
Resources
PoLAR teaches robots to represent motion in a smarter geometric space, helping them better separate how far a change is from what kind of change it is.
Key results
Large-scale pretraining dataset used for the discrete tokenizer and Prismatic-7B VLA.
Average success with joint fine-tuning, compared with 37.0% for the flat tokenizer.
Average success across four simulated manipulation tasks.
Overall final-task success across three real-robot tasks.
Hyperbolic PoLAR score, versus 0.160 for Euclidean PoLAR.
What the paper found
Researchers at Seoul National University introduce PoLAR, or Polar Latent Actions with Radial structure, to separate transition extent from transition mode in robot latent actions. Instead of forcing one code to represent both how far a state changes and what kind of change occurs, PoLAR uses temporal offset as weak ordinal supervision: larger observation gaps are pushed to larger radii, while directions encode mode. Its latent space is implemented in hyperbolic geometry, where angular capacity expands with radius, and training combines start-conditioned ordering and origin-centered radial losses. For discrete actions, a factorized codebook represents each transition with one radial token and four direction tokens. In large-scale pretraining, the method uses frozen DINOv2 features, BridgeData V2’s 60,096 demonstrations, and a Prismatic-7B VLA. PoLAR reaches 70.2 percent average success with joint fine-tuning on RoboMimic and MimicGen, versus 37.0 percent for a flat tokenizer, and 62.5 percent on SimplerEnv-WidowX, exceeding UniVLA at 46.9 percent and π0.5 at 56.2 percent. On three real-world WidowX tasks, it achieves a 76.7 percent overall final-task average, outperforming SmolVLA, UniVLA, Villa-X, and π0.5. Diagnostics show that hyperbolic PoLAR’s latent actions explain 0.184 of ground-truth action variance, compared with 0.160 for its Euclidean variant, supporting the claim that latent-action geometry improves transfer from visual pretraining to robot control.
Original abstract
Latent action pretraining learns representations of visual change from pairs of observations, but existing methods typically encode each transition as a single unstructured representation that entangles transition extent and transition mode. We introduce Polar Latent Actions with Radial structure (PoLAR), which imposes a radial-direction structure on latent actions, encouraging radius to encode transition extent and direction to retain transition mode. PoLAR uses temporal offset between two observations as a weak proxy for transition extent, encouraging latent action from observation pairs separated by larger temporal gaps to occupy larger radii. We instantiate this structure in hyperbolic space, whose expanding volume with radius offers a natural fit for more diverse transition modes at larger extents. Across in-task and large-scale pretraining settings, PoLAR improves downstream policy performance in simulation and real-world robot experiments, outperforming latent action baselines and strong pretrained VLAs. These results suggest that the geometry of the latent action space is an important design choice for transferring visual pretraining to downstream robot policy learning.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.