NTH

SlotDiT: Object-Centric Representations for Diffusion Transformers

AuthorsGjergj Plepi, Sven Behnke

September 18, 2026 2 min read
Watch on YouTube
The one-line take

SlotDiT uses object-centric slots instead of conventional image latents to make diffusion-based robotic video prediction more structured, efficient, and task-effective.

Key results

87.8%
CLIPort video task success

VLM-judged success rate for SlotDiT text-guided video generation.

79.7%
LanguageTable-Synthetic video task success

VLM-judged success rate for SlotDiT text-guided video generation.

73.0%
CLIPort robot-control success

Open-loop pick-and-place success rate for SlotDiT.

9.33
SlotDiT throughput

Frames per second when generating predictions on CLIPort.

5.68×
Speedup over DiT + SD-VAE

Inference speedup measured on an NVIDIA A6000.

10
SlotDiT tokens per frame

Compact slot representation, compared with 256 tokens per frame for VAE-based DiT baselines.

What the paper found

SlotDiT replaces the dense VAE latents used by models such as Stable Diffusion with compact object-centric slots: a frozen DINOv2-ViT and Slot Attention decompose a reference image into entity-level tokens, then a text-conditioned Diffusion Transformer autoregressively denoises future slot trajectories using T5-small text embeddings, adaptive layer normalization, v-prediction, classifier-free guidance, and DDIM sampling. In a controlled comparison against SD-VAE, ImageVAE, VA-VAE, VideoVAE, DINOv2 feature latents, and the autoregressive TextOCVP, SlotDiT prioritizes task-relevant structure over pixel fidelity across CLIPort, LanguageTable-Synthetic, LanguageTable-Real, and Bridgev2. It reaches 87.8% video task success on CLIPort and 79.7% on LanguageTable-Synthetic, while its robot-control success reaches 73.0% on CLIPort, far above VAE-based DiT baselines. The trade-off is weaker perceptual metrics in some settings, showing that visual quality does not reliably predict instruction following. Efficiency is the clearest systems gain: SlotDiT processes 10 tokens per frame instead of 256, generates 9.33 FPS, and achieves a 5.68× speedup over DiT with SD-VAE on an NVIDIA A6000. The study concludes that object-centric structure is a useful inductive bias for diffusion-based robotic world modeling, while noting fixed slot counts, reduced appearance detail, and limited real-robot control evaluation as open issues.

Original abstract

Text-conditioned latent diffusion models perform strongly in video generation and are promising backbones for robotic applications. However, existing approaches rely on pixel-level or VAE-based latent representations that lack explicit semantic structure, leaving the impact of the representation space largely unexplored. Slot-based object-centric representations offer a structured alternative by decomposing scenes into object-level latents, or slots. While they have shown success in dynamics modeling and planning, they have not yet been explored for diffusion-based generative modeling. We introduce SlotDiT, a text-guided Diffusion Transformer (DiT) that operates in a slot-based latent space. Given a reference image and a language instruction, SlotDiT decomposes the scene into object-centric slots representing individual entities. Conditioned on the instruction and observed scene context, the model autoregressively denoises future slot trajectories to predict scene dynamics. To systematically investigate latent-space design for diffusion transformers, we compare slot-based representations against VAE-based and semantics-aligned alternatives within a unified DiT framework. Our experiments show that using slots as DiT latents yields competitive video generation quality while consistently improving task-completion rates across four robotic datasets. Furthermore, their compact representation provides a computationally efficient alternative to VAE-based and semantics-aligned latent spaces. Overall, our results demonstrate that object-centric structure is a powerful inductive bias for diffusion-based generative modeling in robotic environments. The project page is available at https://slot-dit.github.io/.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis