NTH

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

AuthorsYoungjoon Jeong, Jihwan Yu, Minsoo Jo, Junha Chun, Taesup Kim

July 21, 2026 2 min read
Watch on YouTube
The one-line take

PoLAR teaches robots to represent motion in a smarter geometric space, helping them better separate how far a change is from what kind of change it is.

Key results

60,096
BridgeData V2 demonstrations

Large-scale pretraining dataset used for the discrete tokenizer and Prismatic-7B VLA.

70.2%
RoboMimic and MimicGen PoLAR success

Average success with joint fine-tuning, compared with 37.0% for the flat tokenizer.

62.5%
SimplerEnv-WidowX PoLAR success

Average success across four simulated manipulation tasks.

76.7%
Real-world WidowX average

Overall final-task success across three real-robot tasks.

0.184
Action-variance probe R2

Hyperbolic PoLAR score, versus 0.160 for Euclidean PoLAR.

What the paper found

Researchers at Seoul National University introduce PoLAR, or Polar Latent Actions with Radial structure, to separate transition extent from transition mode in robot latent actions. Instead of forcing one code to represent both how far a state changes and what kind of change occurs, PoLAR uses temporal offset as weak ordinal supervision: larger observation gaps are pushed to larger radii, while directions encode mode. Its latent space is implemented in hyperbolic geometry, where angular capacity expands with radius, and training combines start-conditioned ordering and origin-centered radial losses. For discrete actions, a factorized codebook represents each transition with one radial token and four direction tokens. In large-scale pretraining, the method uses frozen DINOv2 features, BridgeData V2’s 60,096 demonstrations, and a Prismatic-7B VLA. PoLAR reaches 70.2 percent average success with joint fine-tuning on RoboMimic and MimicGen, versus 37.0 percent for a flat tokenizer, and 62.5 percent on SimplerEnv-WidowX, exceeding UniVLA at 46.9 percent and π0.5 at 56.2 percent. On three real-world WidowX tasks, it achieves a 76.7 percent overall final-task average, outperforming SmolVLA, UniVLA, Villa-X, and π0.5. Diagnostics show that hyperbolic PoLAR’s latent actions explain 0.184 of ground-truth action variance, compared with 0.160 for its Euclidean variant, supporting the claim that latent-action geometry improves transfer from visual pretraining to robot control.

Original abstract

Latent action pretraining learns representations of visual change from pairs of observations, but existing methods typically encode each transition as a single unstructured representation that entangles transition extent and transition mode. We introduce Polar Latent Actions with Radial structure (PoLAR), which imposes a radial-direction structure on latent actions, encouraging radius to encode transition extent and direction to retain transition mode. PoLAR uses temporal offset between two observations as a weak proxy for transition extent, encouraging latent action from observation pairs separated by larger temporal gaps to occupy larger radii. We instantiate this structure in hyperbolic space, whose expanding volume with radius offers a natural fit for more diverse transition modes at larger extents. Across in-task and large-scale pretraining settings, PoLAR improves downstream policy performance in simulation and real-world robot experiments, outperforming latent action baselines and strong pretrained VLAs. These results suggest that the geometry of the latent action space is an important design choice for transferring visual pretraining to downstream robot policy learning.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis