NTH

Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain

AuthorsYuan Yao, Jin Song, Huixia Li, Tongtong Yuan, Jiaqi Wu, Yu Zhang

June 17, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that even pure noise can help learn better models in semi-supervised transfer, if it is used as a clever surrogate source domain.

Key results

67.90
CIFAR-10 top-1 accuracy

NAF with ResNet-18 on CIFAR-10

55.55
CIFAR-10 ERM top-1 accuracy

Baseline ERM with ResNet-18 on CIFAR-10

49.04
CIFAR-100 top-1 accuracy

NAF with ResNet-18 on CIFAR-100

41.43
CIFAR-100 ERM top-1 accuracy

Baseline ERM with ResNet-18 on CIFAR-100

37.10
ImageNet-1K top-1 accuracy

NAF with ResNet-18 on ImageNet-1K

36.11
ImageNet-1K ERM top-1 accuracy

Baseline ERM with ResNet-18 on ImageNet-1K

What the paper found

Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain introduces a setting where a synthetic noise domain, generated from random distributions such as Gaussian noise, is used as a surrogate source domain for target classification when only a few target labels are available. The core idea is that noise need not be semantically meaningful to help learning: if noise classes are assigned a one-to-one mapping to target class indices, then training a shared classifier on both domains can induce a discriminative class structure in the noise space and transfer that structure to the target space. The paper proves a generalization bound for SSNA and turns it into the Noise Adaptation Framework, which jointly minimizes labeled target risk, noise risk, and a distribution-alignment term in a shared representation space; its preferred alignment objective is Negative Domain Similarity. On CIFAR-10 with ResNet-18, NAF reaches 67.90% accuracy versus 55.55% for ERM, a 12.35 percentage point gain, and on CIFAR-100 it improves from 41.43% to 49.04%. It also scales to fine-grained recognition, reaching 50.86% on CUB-200, 86.58% on OxfordFlowers-102, and 35.75% on StanfordCars-196, while ImageNet-1K improves from 36.11% to 37.10%. The method plugs into SSL systems like UDA and FixMatch, adding up to 20.99 points on CIFAR-10 at epoch 20 for UDA + NAF. Ablations show that both noise risk and alignment matter, that cosine-based alignment beats Euclidean distance, and that removing the class-discriminative structure of noise collapses performance, confirming that the transferable signal comes from structured noise rather than semantics.

Original abstract

Transfer learning aims to facilitate the learning of a target domain by transferring knowledge from a source domain. The source domain typically contains semantically meaningful samples (*e.g.*, images) to facilitate effective knowledge transfer. However, a recent study observes that the noise domain constructed from simple distributions (*e.g.*, Gaussian distributions) can serve as a surrogate source domain in the semi-supervised setting, where only a small proportion of target samples are labeled while most remain unlabeled. Based on this surprising observation, we formulate a novel problem termed *Semi-Supervised Noise Adaptation* (SSNA), which aims to leverage a synthetic noise domain to improve the generalization of the target domain. To address this problem, we first establish a generalization bound characterizing the effect of the noise domain on generalization, based on which we propose a Noise Adaptation Framework (NAF). Extensive experiments demonstrate that NAF effectively leverages the noise domain to tighten the generalization bound of the target domain, leading to improved performance. The codes are available at https://github.com/AIResearch-Group/SSNA.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis