JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Follow research on robot perception, control, and learning. Explore how experimental results translate across tasks, environments, and physical systems.
50 papers · Latest edition October 4, 2026
Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.
Newest editions first.
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.
Jake Gonzales, Arturo Flores Alvarez, Yu-Ming Chen, Aaron D. Ames, Lillian J. Ratliff, Manikantan Nambi
LIMBO teaches humanoid robots to discover and internalize safety boundaries so they can perform agile maneuvers without relying on an online safety filter.
Yujie Xiong, Peng Zhai, Taixian Hou, Quancheng Qian, Cunwang Liu, Kangmai Hu, Long Yang, Zhiyan Dong, Lihua Zhang
SwingBot teaches humanoid robots to swing continuously across overhead bars by combining structured motion guidance with learned sensing and control.
Jie Yin, Wanli Xing, Zeyuan Zhao, Xuezhou Zhu, Zhijie Deng, Kaifeng Zhang
TacBPM gives dexterous robots a tactile-aware library of reusable motion skills so they can reorient unfamiliar objects more reliably and with less trial-and-error.
Ziqin Huang, Yingyue Li, Chenyangguang Zhang, Ruida Zhang, Yuxin Chen, Gu Wang, Xingyu Liu, Masayoshi Tomizuka, Xiangyang Ji
3DWay helps robots plan more reliably by turning multi-view visual predictions into geometrically consistent 3D movement waypoints.
Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang
AdaRoboVLG lets robots adapt how they grasp objects by combining a reusable physical grasping policy with task-specific knowledge from foundation models.
Kai Stewart, Yasunori Toshimitsu, Robert K. Katzschmann
A robot hand learns to write with a grasped pen in seconds by estimating how its movements affect the pen in real time, without simulation or demonstrations.
Seungyeon Kim, Noémie Jaquier
ChainSplat combines physics-inspired articulated modeling with Gaussian splatting to reconstruct and control deformable objects such as cables from ordinary multi-view video.
Satvik Sharma, Samrat Sahoo, Huang Huang, Fei-Fei Li Jiajun Wu, Dorsa Sadigh, Jeannette Bohg
DemoMimic teaches robot hands to generalize dexterous manipulation by concentrating on the local geometry where contact occurs.
Taeyoon Lee, Chunpeng Wang, Christopher G. Atkeson, Alfred A. Rizzi, Nicolas Rojas
A bimanual robot learns five juggling patterns safely in under five minutes by refining imperfect prior models with its own real-world experience.
Julien Merand, Boris Meden, Liming Chen, Mathieu Grossard
CoToGrasp learns to generate functionally meaningful dexterous grasps across unseen objects by conditioning on contact topology rather than object-specific annotations.
Chenhui Pan, Tong Xu, Francesco Cancelliere, Xuesu Xiao
NeSAM helps off-road robots predict and navigate deformable terrain by blending soil physics with learned dynamics that adapt online.
Ruihua Han, Rui Gao, Zhe Liu, Xinyi Wang, Chang Chen, Shuai Wang, Qi Hao, Jia Pan, Hengshuang Zhao
SRL-MPC blends reinforcement learning with safety-constrained model predictive control to help differently shaped robots navigate dense crowds more safely and efficiently.
Zili Tang, Tiecheng Guo, Qinyue Zhang, Meng Guo
FlexWorm helps soft, suction-powered robots navigate complex surfaces by combining geometric planning with learned motion snippets and fallback search.
Yijie Xu, Haopeng Jin, Run Zhou, Shengbang Liu, Sixiang Chen, Hongyang Cheng, Sicheng Hu, Peterson Co, Jinwen Luo, Huajie Tan, Shanghang Zhang
Robo-Dopamine 2.0 gives robot policies a history-aware, OOD-sensitive reward signal to recover from mistakes and improve manipulation reliability.
XPolicyLab Community, Tianxing Chen, Yue Chen, Tian Nian, Zijian Cai, Guangyu Chen, Wenwei Lin, Qiwei Liang, Peicheng Xiang, Kailun Su, Zixuan Li, Junyuan Tang, Yan Qin, Qiangyu Chen, Shaolong Zhu, Xiang Li, Jiahao Zhang, Weijie Wan, Baijun Chen, Honghao Su, Kehe Ye, Shujia Liu, Kaixuan Wang, Haotian Liang, Yunze Liu, Mingleyang Li, Yuran Wang, Boyu Chen, Hongzhe Bi, Shuhe Huang, Hengkai Tan, Jisong Cai, Yao Mu, Jun Guo, Xiaofeng Wang, Zheng Zhu, Weijie Ke, Hengtao Li, Yuhang Tang, Xiaofan Li, Ganlin Yang, Zhangzheng Tu, Shuai Yang, Wenxuan Song, Pengxiang Ding, Kaidong Zhang, Yu Sun, Junliang Guo, Tong Zhang, Yixing Chen, Rongxu Cui, Zongzheng Zhang, Haoxiang Ma, Junhao Cai, Haoyu Zhang, Senqiao Yang, Jinhui Ye, Pengguang Chen, Shu Liu, Xiu Su, Wenhan Fang, Wenhao Li, Yichao Cao, Chengyao Wang, Qiang Chen, Ping Luo, Wenbo Ding
XPolicyLab aims to make evaluating and deploying robot policies as plug-and-play as using a common software standard instead of building a new integration for every policy and environment.
Giovanbattista Gravina, Luca Rossini, Carlo Rizzardo, Arturo Laurenzi, Nikos Tsagarakis
A reinforcement-learning controller helps a large quadruped keep walking after actuator failures by learning when to change its gait timing.
Renhao Lu, Mingxin Wang, Chenyang Cao, Yang Yang, Guoping Pan, Kangkang Dong, Yi Cheng, Houde Liu
Push-Wiper teaches robots to gather rather than spread messy stains, enabling more general cleaning across materials, residue types, and curved surfaces.
Dylan Miller, Martin Jagersand
Temporal Policy makes diffusion-like robot control faster by starting action generation from the robot’s recent history instead of random noise.
Lifeng Zhuo, Wendi Chen, Han Xue, Shirun Tang, Jun Lv, Cewu Lu, Chuan Wen
FA-RDP lets robot policies explore diverse approaches before contact and react quickly to force feedback once manipulation becomes constrained.
Xiaofan Lu, Kaiji Huang, Jiahui Chen, Yuankai Lin, Hua Yang, Zhouping Yin
FasTac is a fast curved tactile fingertip that uses multispectral vision and FPGA processing to sense 3D contact shape and forces during dexterous robot manipulation.
Filippo Lazzati, Kyle Stachowicz, William Chen, Alberto Maria Metelli, Andrew Wagenmaker, Sergey Levine
Action chunking helps robots not just by planning farther ahead, but by implicitly ensembling different temporal views of the task.
Mengfei Zhao, Dihong Huang, Yikai Tang, Peihao Li, Mingxuan Yan, Ruiqi Zhuang, Yanjia Huang, Jie Wang, Hai Zhai, Tony Zhou, Rui Zhang, Zhexi Luo, Yuchen Huang, Jianfei Yang, Jiachen Li
AXIS turns community-collected robot demonstrations into a scalable data engine that measurably improves vision-language-action manipulation policies.
Panagiotis Mermigkas, Argyris Manetas, Petros Maragos
GLAM-SLAM makes Gaussian-splatting-based robot mapping faster and more scalable for large outdoor environments.
Wun Lam Yeung, Wenjun Liu, Yui Cheung Yu, Zhengyan Lambo Qin, Qijin She, Heng Li, Ziqi Wang, Ping Tan
A diffusion-and-energy-based robot planner helps two arms grasp, hand over, rotate, and place objects more reliably and smoothly.
Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu
ABot-N1 is a navigation foundation model that turns language and vision into pixel-level goals before executing actions, improving robustness and interpretability in complex indoor and urban navigation.
Yifan Zhong, Zhang Chen, Tianrui Guan, Fanlian Zeng, Yuyao Ye, Tianjia He, Ka Nam Lui, Jiayi Li, Tingrui Zhang, Ruilin Yan, Xinhao Ji, Guangyu Zhao, Wenjie Lou, Jiayuan Zhang, Yuanpei Chen, Yaodong Yang
EgoSteer is a full-stack robot learning pipeline that turns massive egocentric human video into a steerable dexterous manipulation policy capable of generalizing to many real-world tasks.
Youngjoon Jeong, Jihwan Yu, Minsoo Jo, Junha Chun, Taesup Kim
PoLAR teaches robots to represent motion in a smarter geometric space, helping them better separate how far a change is from what kind of change it is.
Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng, Fengyuan Hu, Yunhao Ge, Jimmy Wu, Tianyuan Dai, Scott Reed, Li Fei-Fei, Yuke Zhu, Linxi "Jim" Fan
RoboTTT gives robot foundation models an 8K-step memory by learning at test time, enabling stronger imitation, adaptation, and long-horizon manipulation.
Dongyoon Hwang, Byungkun Lee, Dongjin Kim, Hyojin Jang, Hoiyeong Jin, Jueun Mun, Minho Park, Hojoon Lee, Hyunseung Kim, Jaegul Choo
3D HAMSTER makes robot planners output depth-aware 3D waypoints instead of flat 2D paths, helping language-guided manipulation work better in real-world and visually shifting environments.
Hongyu Qu, Jianzhe Gao, Xiaobin Hu, Shaohuan Yang, Xinlei Yu, Rui Yan, Wenguan Wang, Xiangbo Shu, Shuicheng Yan
This paper teaches robot vision-language-action models to remember past experience in a shared latent space, helping them handle longer and more complex manipulation tasks.
Hanan Gani, Tejal Kulkarni, Madhoolika Chodavarapu, Nicklas Hansen, Manmohan Chandraker
RoboTALES teaches robot policies by using language-guided planning and vision-language feedback to imagine task-aligned futures, improving long-horizon manipulation.
Kinam Kim, Namiko Saito, Heecheol Kim, Katsushi Ikeuchi, Jaegul Choo, Yasuyuki Matsushita
This paper shows how a small simulation-trained corrective policy can make vision-language-action robot policies much more reliable in the real world without extra robot training.
Tyler Ga Wei Lum, Kushal Kedia, C. Karen Liu, Jeannette Bohg
This work shows that teaching a dexterous robot to 'play' first can make it far better at precise real-world assembly tasks like tight insertions and screwing.
Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu, Xiong-Hui Chen
Qwen-RobotManip shows that aligning diverse robot and human demonstration data at scale can unlock stronger generalization for vision-language-action robots across many platforms and real-world settings.
Nadun Ranawaka, Josiah Wong, Wei-Lin Pai, Wei-Teng Chu, Tianyuan Dai, Masoud Moghani, Hang Yin, Yunfan Jiang, Wesley Durbano, Brandon Huynh, Yu Fang, Linxi Fan, Danfei Xu, Ruohan Zhang, Li Fei-Fei, Bowen Wen, Ajay Mandlekar, Yuke Zhu
SimFoundry turns a single real video into editable simulation twins that help robot policies train, generalize, and transfer to the real world more reliably.
Kunyun Wang, Yuhang Zheng, Yupeng Zheng, Jieru Zhao, Wenchao Ding
The paper makes robot control smoother and more continuous by learning fast action chunks in latent space and refining them on the fly for real-world contact tasks.
Ziang Li, Dongzhou Cheng, Yibin Wang, Shiyue Wang, Xiaoyang Xu, Lingxuan Weng, Juan Wang, Jiaqi Wang
Light-WAM makes robot world-action models much lighter by predicting actions from fused backbone states, keeping strong manipulation performance while cutting latency and memory.
Josef Chen
AEGIS helps robots avoid failure spirals by predicting trouble early and handing control to a stronger policy only at the risky moments.
Arman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen, Weiwei Chen, Xuan Zhang, Geng Yuan, Yanzhi Wang
Flash-WAM makes diffusion-based robot world models fast enough for real-time control by using different distillation parameterizations for video and action streams, cutting inference from seconds to milliseconds without losing much task success.
Barak Or
This paper asks a simple but important question: can a model’s predicted action or dynamics actually work in the real physical world, and it proposes a gate to reject implausible proposals before execution.
Tongyan Fang, Siyuan Huang, Naiyu Fang, Ganlong Zhao, Zhongjin Luo, Jianbo Liu, Xiaogang Wang, Ying Dong, Hongsheng Li
This paper makes robot fine-tuning smarter by turning sparse success/failure signals into more informative, stage-aware training weights, boosting real-world task success on bimanual manipulation.
Junyi Zhang, Jiaxin Ge, Hanjun Yoo, Letian Fu, Zihan Yang, Yaowei Liu, Raj Saravanan, Shaofeng Yin, Justin Yu, Dantong Niu, Zirui Wang, Roei Herzig, Ken Goldberg, Yutong Bai, David M. Chan, Ion Stoica, Angjoo Kanazawa, Jiahui Lei, Haiwen Feng, Trevor Darrell
This paper teaches robots to play first, so they can build reusable code-based skills that help them solve future tasks better.
Xingyao Lin, Guojin Zhong, Tianyi Lu, Ziyi Ye, Yichen Zhu, Zuxuan Wu, Yu-Gang Jiang
ActiveMimic teaches robots from human first-person videos by treating camera motion as useful action, helping bridge the gap between egocentric video and robot pretraining.
Jianwei Tai
This paper shows that in robot vision-language-action systems, the same checkpoint can behave like a different policy once action normalization metadata changes, creating a hidden safety risk before rollout.
Rui Zhao, Kaiming Yang, Jifeng Zhu, Siyang Chen, Ziqi Wang, Weijia Wu, Kevin Qinghong Lin, Heng Wang, Mike Zheng Shou
Dream.exe tests whether video generators can dream motions that actually work in a robot simulator, turning pretty videos into a practical measure of physical understanding.
Zehao Yu, Jiakun Zheng, Weiji Xie, Jiyuan Shi, Chenyun Zhang, Chenjia Bai, Xuelong Li
OASIS shows that carefully built simulated data can sometimes beat real teleoperation for humanoid robot loco-manipulation on the physical robot.
Christian Bianchi, Siamak Yousefi, Alessio Sampieri, Andrea Roberti, Luca Rigazio, Fabio Galasso, Luca Franco
WIZARD teaches a robot policy to adapt to new tasks in one shot by predicting the right weight updates from a demo video and instruction, avoiding task-specific fine-tuning.