TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation
AuthorsAdeela Islam, Zorah Lähner, Vittorio Murino, Vladislav Golyanik
Resources
TokenMatch uses curvature-aware mesh tokens and attention to rapidly match corresponding points across difficult 3D shapes, including partial and heavily deformed ones.
Key results
Mean intersection-over-union score for partial-to-partial matching.
Mean intersection-over-union score under non-isometric animal deformations.
Mean intersection-over-union score on the challenging partial-shape benchmark.
Mean geodesic error for full-shape matching without retraining.
Seconds per shape pair for overlap, functional-map, and point-match recovery.
Number of geometry-aware tokens processed for each shape.
What the paper found
TokenMatch tackles 3D mesh correspondence when shapes are partially observed or undergo strong non-isometric deformation. Its central innovation is curvature-guided tokenisation: curvature-weighted farthest-point sampling combines local mean curvature with mid-frequency spectral energy, then represents each mesh with overlapping patches assigned through geodesic Gaussian weights. A ViT-Base transformer processes 256 mesh tokens using self-attention for within-shape structure and bidirectional cross-attention for direct inter-shape reasoning. Masked autoencoder pre-training, PointInfoNCE feature supervision, overlap prediction, and 50-by-50 functional maps jointly produce dense point correspondences. Trained exclusively on partial-to-partial BeCoS data, TokenMatch achieves mIoU scores of 85.56 on CP2P24, 85.21 on PSMAL, and 65.25 on BeCoS, outperforming descriptor-based alternatives including DINOv2 features and EchoMatch. Without retraining, it also reaches a mean geodesic error of 3.45 on SHREC’19 and transfers from partial training to full-shape matching. Inference takes 0.16 seconds per shape pair on two NVIDIA A100 80GB GPUs, making the feed-forward model substantially faster than optimization-heavy correspondence methods while remaining robust to irregular mesh sampling and moderate noise.
Original abstract
While data-driven 3D shape correspondence estimation has recently seen substantial progress, robust matching under partial observations and strong non-isometric deformations remains challenging. Existing learning-based approaches often rely on hand-crafted descriptors or template-based representations, whereas recent generative models over functional maps suffer from high inference cost, limited interpretability, and poor generalisation to partial shapes. In response to these limitations, this paper introduces TokenMatch, a new transformer-based unified model for estimating 3D shape correspondences. Our feed-forward approach trained exclusively on BeCoS, a challenging non-isometric partial-to-partial shape-matching dataset, can generalise to matching full shapes without retraining or fine-tuning. TokenMatch uses self- and cross-attention mechanisms to efficiently learn patch-level and point-level relations as well as dense correspondences between shape pairs. Our core insight is that meshes can be adaptively tokenised into patches using shape curvature guidance, enabling effective learning of shape-specific geometric descriptors for correspondence estimation. We evaluate TokenMatch on standard benchmarks for partial and full shape matching, including CP2P, PSMAL, BeCoS, FAUST, SCAPE, and SHREC'19. Our method achieves consistently high performance, in most cases outperforming existing methods for partial and full shape matching in the mean geodesic error and intersection-over-union metrics, while also running faster at sub-second inference speeds.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.