SiamJEPA: On the Role of Siamese Student Encoders in JEPA
AuthorsMakoto Yamada
Resources
SiamJEPA shows that giving JEPA a Siamese student encoder can act like a useful regularizer, improving early learning and representation quality in self-supervised vision models.
Key results
Linear probing accuracy on ImageNet-1K with ViT-Base
MAE baseline under the same 400-epoch evaluation setting
Context Autoencoder comparison point on ImageNet linear probing
Reported original I-JEPA result used as a non-direct comparison
SiamJEPA linear probing accuracy with weak Siamese regularization
What the paper found
SiamJEPA, from Makoto Yamada at the Okinawa Institute of Science and Technology, studies a simple but underexplored question in JEPA-style self-supervised vision learning: what happens if the student network uses Siamese encoders instead of a single encoder. The proposed model combines two independently masked ViT-Base student branches with an exponential moving average teacher, a linear alignment predictor, and a probabilistic latent predictor inspired by PhiNet and stochastic frame prediction. Training optimizes a masked latent prediction term plus a KL regularizer between a posterior inferred from both views and a prior inferred from one view, with disjoint masks used to prevent shortcut learning. On ImageNet linear probing, SiamJEPA reaches 70.2% top-1 accuracy after 400 epochs, compared with 61.9% for MAE at the same 400-epoch budget, while remaining close to CAE’s 70.4% reached only after 1600 epochs; the original I-JEPA baseline is reported at 72.9% after 600 epochs, but the setups are not directly comparable. The ablations show that the Siamese KL term acts mainly as a regularizer and accelerates early convergence: with KL weight 0.01, top-1 accuracy rises to 63.72% at epoch 100 and 69.30% at epoch 300, versus 57.89% and 66.78% for a much weaker 0.0001 setting. Block masking is consistently stronger than random masking, and larger weight decay improves the 400-epoch result to 70.15%, reinforcing the paper’s core claim that Siamese student encoders are not just an architectural variant but a useful inductive bias for predictive representation learning.
Original abstract
Recently, Joint Embedding Predictive Architectures (JEPAs) have attracted significant attention in the computer vision and machine learning communities as a promising framework for self-supervised representation learning. Unlike masked autoencoders that reconstruct pixels, JEPA models learn representations by predicting latent embeddings of masked regions. Existing JEPA-based methods, such as I-JEPA and V-JEPA, typically employ a single encoder in the student network. In contrast, using Siamese encoders for student network is more naturally aligned with brain-inspired representation learning frameworks, yet their role in JEPA models remains largely unexplored. In this paper, we investigate the effect of Siamese student encoders in JEPA-based representation learning. To this end, we propose SiamJEPA, masked Siamese student encoders equipped with an exponential moving average (EMA) teacher network. SiamJEPA can also be viewed as a JEPA formulation of the brain-inspired representation learning model PhiNet. Through extensive experiments on ImageNet linear probing, we demonstrate that Siamese encoders act as an effective regularizer for the JEPA objective, improving representation separability and accelerating learning during the early stages of training. Furthermore, SiamJEPA consistently outperforms comparable single-encoder JEPA variants under limited training budgets and achieves higher linear probing accuracy than Masked Autoencoders (MAE) which requires longer training. Our findings reveal that Siamese student encoders are not merely an architectural choice but constitute an important inductive bias for predictive representation learning. These results provide new insights into the design of JEPA-based models and suggest that incorporating Siamese student architectures offers a simple yet effective approach for improving self-supervised representation learning.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.