Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation
AuthorsChunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li
This paper teaches ultrasound models to learn from anatomy, not just pixels, by using self-supervision tied to clinically meaningful structures for more robust medical imaging representations.
Key results
AnaUS reaches 96.2% accuracy on the POCUS classification benchmark, outperforming the strongest ultrasound-specific baseline IVPP++ by 3.0 points.
AnaUS reaches 87.1% accuracy on BUSI classification, beating IVPP++ by 3.6 points.
AnaUS achieves 85.9% Dice on the UDIAT-B segmentation benchmark, exceeding IVPP++ by 2.3 points.
Removing LP-SAM causes a 15.7-point Dice drop on UDIAT-B, showing anatomy-aware mask generation is the dominant contributor.
The retained ResNet-18 encoder processes a 224×224 image in about 0.45 ms on an RTX 4090.
What the paper found
Anatomy-Anchored Self-Supervision, or AnaUS, reframes ultrasound pre-training around clinically meaningful anatomical structures instead of whole frames or generic image regions. The key novelty is LP-SAM, a two-stage adaptation of Segment Anything Model using lightweight adapters and a learnable latent prompt engine with Cross-Perception Attention, which turns public image-mask pairs from DDTI, TG3K, and CAMUS into annotation-free anatomical masks at scale. AnaUS then combines two self-supervised objectives on unlabeled Butterfly and CAMUS data: multi-scale, BYOL-style anatomy contrast that pulls together the same structure across views while repelling different structures, and a contextual prediction loss that reconstructs corrupted core regions using ultrasound-specific perturbations such as pixel shuffling, Gaussian speckle noise, and occlusion. Across six public benchmarks, including POCUS, BUSI, UDIAT-B, TN3K, and HMC-QU, AnaUS reaches 96.2 percent accuracy on POCUS, 87.1 percent on BUSI, and 85.9 percent Dice on UDIAT-B, outperforming the strongest ultrasound-specific baseline IVPP++ by 3.0 points, 3.6 points, and 2.3 points respectively. Ablations show that removing LP-SAM drops UDIAT-B Dice by 15.7 points, making anatomy-aware mask generation the dominant factor behind the gain. The retained ResNet-18 encoder remains real-time, processing a 224×224 image in about 0.45 milliseconds on an RTX 4090.
Original abstract
Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning. In this work, we propose an anatomy-anchored ultrasound self-supervision framework ANAUS that shifts representation learning from generic visual regions to clinically meaningful anatomical structures. Utilizing a learnable latent prompt engine alongside a one-time domain adaptation on existing public image--mask pairs, we empower the LP-SAM module to achieve annotation-free anatomy delineation at scale. Building upon this anatomical grounding, we propose a dual-policy self-supervised learning paradigm consisting of inter-view semantics-aware anatomy-separating alignment and contextual core-region prediction to enhance representation learning. Specifically, the former enforces feature invariance within identical anatomical regions while promoting discriminability across distinct structures; the latter compels the model to reconstruct corrupted regions, thereby capturing fine-grained structural details. Extensive evaluations on six public datasets demonstrate that \ours{} consistently outstrips current state-of-the-art methods while maintaining the computational efficiency essential for clinical deployment. Code is available at https://github.com/zhcz328/ANAUS.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.