U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training
AuthorsZhiwen Yang, Jiayin Li, Hao Lu, Hui Zhang, Zihua Wang, Bingzheng Wei, Yan Xu
U-TTT adapts a PET denoising model on the fly during inference, using self-supervised spatial and frequency updates to stay robust across scanners and dose shifts.
Key results
U-TTT model size
U-TTT computational cost
Average PSNR improvement over the strongest baseline on in-distribution data
Best PSNR at DRF=2 on the base dataset
Average PSNR on unseen dose levels
Average PSNR on unseen scanners
What the paper found
U-TTT, developed by researchers at Beihang University, Tsinghua University, and ByteDance, is a 3D PET denoising framework that tackles the core failure mode of fixed-parameter models under scanner and dose distribution shift by embedding test-time training directly into a U-shaped network. The method replaces standard attention with two differentiable adaptation blocks: S-TTT for spatial structural correction and F-TTT for frequency-domain noise suppression, each performing online self-supervised reconstruction updates during inference so the model adapts to each test volume before predicting the restored full-dose PET image. On a base dataset with DRF=2, 3, 6, and 12, U-TTT reaches 51.63 dB, 50.01 dB, 48.02 dB, and 45.96 dB PSNR with 0.9823, 0.9741, 0.9621, and 0.9498 SSIM, while reducing lesion error to 0.1486 on average and staying lightweight at 10.20M parameters and 43.52 GFLOPs. Against five PET denoising baselines, it improves average PSNR by 0.80 dB over the strongest competitor on in-distribution data and also generalizes best on out-of-distribution tests, reaching 46.86 dB on unseen dose settings and 43.10 dB on unseen scanners. Ablations show that combining S-TTT and F-TTT is necessary, and the proposed gated inner model outperforms linear and MLP alternatives, confirming that dual-domain test-time adaptation is the key contribution.
Original abstract
Existing deep learning models for Positron Emission Tomography (PET) image denoising often suffer from severe performance degradation under distribution shifts, fundamentally restricting their robust clinical deployment. This lack of generalization stems from the conventional paradigm of fixed-parameter models that cannot adapt to variations in test data (e.g., dose levels or scanner types) after training. To overcome this limitation and achieve robust generalization, we introduce U-TTT, a novel U-shaped model that integrates Test-Time Training (TTT) layers to dynamically adjust model parameters during inference through self-supervision, thereby adapting to the specific characteristics of each test instance. Furthermore, to comprehensively capture the complex degradations of 3D PET data, U-TTT features a dual-domain adaptation mechanism comprising a Spatial Test-Time Training (S-TTT) layer and a Frequency Test-Time Training (F-TTT) layer. The S-TTT layer captures and corrects spatial structural degradations, while the F-TTT layer suppresses global noise spectra and restores delicate high-frequency details. Extensive experiments demonstrate that U-TTT achieves state-of-the-art PET denoising performance and exhibits superior generalization under challenging distribution shifts, including both unseen dose levels and unseen scanners. Our code will be available at https://github.com/Yaziwel/U-TTT.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.