NTH

MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception

AuthorsYuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu

AffiliationsWave Intelligence Lab · Department of Computer Science and Engineering, The Chinese University of Hong Kong · Shanghai Artificial Intelligence Laboratory* liyh5366@gmail.com

October 8, 2026 2 min read
Watch on YouTube
The one-line take

MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.

Key results

17
Single-trajectory best evaluations

Best result in 17 of 30 MISATO single-trajectory evaluations.

26
Single-trajectory top-two evaluations

Ranks among the top two in 26 of 30 MISATO evaluations.

9
Multi-trajectory evaluations won

Best performance across all 9 evaluations on MISATO-100, MISATO-1000, and MISATO-Full.

35.66
KTH MSE

Protocol-matched best-run MSE for KTH video prediction.

9.9%
KTH MSE improvement

MSE reduction relative to the strongest baseline.

2.5M
KTH parameter count

Model size for the KTH video-prediction system.

What the paper found

MinkowskiPE introduces a relative positional encoding for spatiotemporal Transformers that treats every token as an event with joint time and space coordinates. It projects each event coordinate into learnable scalar directions, then applies Lorentz boosts to two-dimensional query and key feature blocks while leaving values unchanged. Because the Lorentz metric converts absolute transformations into a function of relative spacetime displacement, attention becomes invariant to global coordinate translation without changing the standard dot-product interface or compatibility with FlashAttention. On the MISATO molecular-dynamics benchmark, MinkowskiPE achieves the best result in 17 of 30 single-trajectory evaluations, ranks in the top two for 26 of 30, and wins all 9 evaluations across MISATO-100, MISATO-1000, and MISATO-Full. On KTH video prediction, it reaches an MSE of 35.66, lowers MSE by 9.9 percent against the strongest baseline, and uses 2.5M parameters. Controlled ablations show a 5.7 percent MSE reduction on KTH when MinkowskiPE replaces no positional encoding. The same geometric mechanism works for molecular coordinates in three dimensions and video grids in two dimensions, while implementation uses FlashAttention on NVIDIA H200 GPUs; the paper also reports ChatGPT and Codex assistance for writing, literature discovery, debugging, and experiment organization.

Original abstract

Modeling spatiotemporal coupling is a key challenge in building physical intelligence across scales, from microscopic to macroscopic. Existing models capture such structure broadly through physics-motivated dynamical formulations or learning-motivated architectures. The former provide stronger priors but may constrain flexibility, whereas the latter are more flexible but leave the spatiotemporal coupling largely implicit. We therefore seek an approach that combines flexible learning with an explicit geometric bias for jointly modeling time and space. To this end, we propose Minkowski Positional Encoding (MinkowskiPE), which uses joint temporal and spatial coordinates to parameterize Lorentz transformations applied to query and key features. With MinkowskiPE, the query-key attention score depends on position only through the relative spacetime displacement between the two tokens and is therefore invariant to global translation of the coordinates. This paradigm retains the standard dot-product attention interface and remains compatible with efficient attention implementations. We evaluate MinkowskiPE on microscopic molecular dynamics and macroscopic video prediction tasks, achieving the best results on all nine multi-trajectory molecular evaluations and reducing KTH video-prediction MSE by 9.9% relative to the best baseline while using roughly one-tenth as many parameters.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis
03Attention

Block Sparse Attention with Log-Linear Complexity

Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu

PISA makes long-context attention more scalable by hierarchically narrowing relevant key blocks instead of comparing every query with every block.

Read analysis