Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain
AuthorsYuan Yao, Jin Song, Huixia Li, Tongtong Yuan, Jiaqi Wu, Yu Zhang
This paper shows that even pure noise can help learn better models in semi-supervised transfer, if it is used as a clever surrogate source domain.
Key results
NAF with ResNet-18 on CIFAR-10
Baseline ERM with ResNet-18 on CIFAR-10
NAF with ResNet-18 on CIFAR-100
Baseline ERM with ResNet-18 on CIFAR-100
NAF with ResNet-18 on ImageNet-1K
Baseline ERM with ResNet-18 on ImageNet-1K
What the paper found
Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain introduces a setting where a synthetic noise domain, generated from random distributions such as Gaussian noise, is used as a surrogate source domain for target classification when only a few target labels are available. The core idea is that noise need not be semantically meaningful to help learning: if noise classes are assigned a one-to-one mapping to target class indices, then training a shared classifier on both domains can induce a discriminative class structure in the noise space and transfer that structure to the target space. The paper proves a generalization bound for SSNA and turns it into the Noise Adaptation Framework, which jointly minimizes labeled target risk, noise risk, and a distribution-alignment term in a shared representation space; its preferred alignment objective is Negative Domain Similarity. On CIFAR-10 with ResNet-18, NAF reaches 67.90% accuracy versus 55.55% for ERM, a 12.35 percentage point gain, and on CIFAR-100 it improves from 41.43% to 49.04%. It also scales to fine-grained recognition, reaching 50.86% on CUB-200, 86.58% on OxfordFlowers-102, and 35.75% on StanfordCars-196, while ImageNet-1K improves from 36.11% to 37.10%. The method plugs into SSL systems like UDA and FixMatch, adding up to 20.99 points on CIFAR-10 at epoch 20 for UDA + NAF. Ablations show that both noise risk and alignment matter, that cosine-based alignment beats Euclidean distance, and that removing the class-discriminative structure of noise collapses performance, confirming that the transferable signal comes from structured noise rather than semantics.
Original abstract
Transfer learning aims to facilitate the learning of a target domain by transferring knowledge from a source domain. The source domain typically contains semantically meaningful samples (*e.g.*, images) to facilitate effective knowledge transfer. However, a recent study observes that the noise domain constructed from simple distributions (*e.g.*, Gaussian distributions) can serve as a surrogate source domain in the semi-supervised setting, where only a small proportion of target samples are labeled while most remain unlabeled. Based on this surprising observation, we formulate a novel problem termed *Semi-Supervised Noise Adaptation* (SSNA), which aims to leverage a synthetic noise domain to improve the generalization of the target domain. To address this problem, we first establish a generalization bound characterizing the effect of the noise domain on generalization, based on which we propose a Noise Adaptation Framework (NAF). Extensive experiments demonstrate that NAF effectively leverages the noise domain to tighten the generalization bound of the target domain, leading to improved performance. The codes are available at https://github.com/AIResearch-Group/SSNA.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.