NTH

Transfer Learning in Nonparametric Regression with Deep ReLU Networks

AuthorsJunpeng Ren, Carlos Misael Madrid Padilla, Yanzhen Chen, Oscar Hernan Madrid Padilla

August 23, 2026 2 min read
Watch on YouTube
The one-line take

A new theory and method lets deep ReLU networks borrow strength across related datasets to improve nonparametric regression.

Key results

23705
UTKFace dataset size

Facial images used for grouped age-regression transfer learning.

512
UTKFace feature dimension

FaceNet latent features extracted from each image.

59.6
UTKFace 2-Stage MSE

Overall age-estimation MSE achieved by the proposed method.

61.8
UTKFace pooled MSE

Overall MSE for pooled neural-network training.

50000
Scenario 1 sample size

Largest simulation sample size reported.

0.27
Scenario 1 2-Stage MSE

MSE ×10^2 for two-stage neural transfer learning at sample size 50000.

What the paper found

This paper introduces a two-stage offset transfer-learning framework for grouped nonparametric regression, assuming each group’s conditional mean is the sum of a shared function and a group-specific deviation. The first stage pools all groups to estimate the overall mean, while the second learns each group’s residual offset with a separate estimator, then adds the two components. General L2 error bounds cover broad nonparametric estimators, growing numbers of groups, sub-exponential noise, sample splitting, trend filtering, and orthogonal series regression. When both stages use dense deep ReLU networks under hierarchical composition models, the resulting convergence rates depend on intrinsic component dimensions rather than ambient dimension, overcoming the curse of dimensionality; positive transfer is strongest when the shared mean or offsets are simpler than the full group functions, or when pooling supplies substantially more data. In simulations over 50 Monte Carlo replications, the two-stage neural estimator consistently outperformed pooled, separate, label-augmented, fine-tuned, random-forest, and ptLasso alternatives; in one additive scenario at sample size 50000, its MSE was 0.27, versus 1.93 for pooled neural regression. On UTKFace, using 512-dimensional features extracted by FaceNet trained on VGGFace2, the method reduced overall age-estimation MSE to 59.6 from 61.8 for pooled training across 23705 images, while achieving the best results for White, Black, and Asian groups.

Original abstract

This paper develops a general transfer learning framework for nonparametric regression with data consisting of multiple groups. Under the assumption that groups share a common structure along with group-specific deviations in additive form, the proposed method employs a two-stage offset learning procedure: the first stage pools data from all groups to estimate an overall mean function, and the second stage estimates offsets for each group, yielding final group-level estimators through additive combination. Upper bounds on the $\mathcal L_2$ error are established for the proposed framework, covering a broad class of nonparametric estimators under mild complexity and noise conditions. When instantiated with deep ReLU networks, explicit convergence rates are derived under hierarchical composition models, demonstrating the ability to overcome the curse of dimensionality. Conditions that enable positive transfer with faster rates are considered, including learning with simpler functions and data augmentation through pooling samples across groups. Various simulations and real-data experiments further validate the effectiveness of the proposed method.

Read the original paper

More in Neural Networks

Browse all 22 papers →
02Neural Network

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.

Read analysis
03Neural Network

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.

Read analysis