Divide and Contrast: Learning Robust Temporal Features without Augmentation
AuthorsAbdul-Kazeem Shamba, Kerstin Bach, Gavin Taylor
This paper introduces a faster, augmentation-free way to learn strong time-series representations by contrasting meaningful chunks of each sequence instead of individual timesteps.
Key results
Di-COT achieves this average accuracy across six large-scale real-world datasets with a frozen backbone and linear probe.
CaTT is the strongest prior temporal-adjacency baseline on the same six-dataset linear-evaluation benchmark, below Di-COT.
TNC is another baseline on the same six-dataset linear-evaluation benchmark, also below Di-COT.
Di-COT has the lowest cumulative training time across the six large-scale datasets among deep learning methods.
In the low-label regime, Di-COT achieves the best average linear-evaluation accuracy across the six datasets.
Di-COT attains the best average clustering normalized mutual information across the six datasets.
What the paper found
Divide and Contrast Learning, or Di-COT, is an unsupervised time-series representation method that removes two major costs in prior self-supervised learning: data augmentation and multiple encoder passes. Instead of contrasting individual timestamps, Di-COT stochastically partitions each window into 2 to 10 overlapping sub-blocks, encodes them with a shared InceptionTime backbone, and treats each sub-block’s immediate predecessor as the positive in a temperature-scaled cross-entropy objective. This reformulation makes the loss independent of sequence length and reduces similarity computation from timestep-level quadratic scaling to an O(Bk^2d) objective, where k is far smaller than T. The design directly addresses false positives caused by assuming every neighboring timestep is semantically similar, a failure mode that hurts methods such as CaTT on heterogeneous benchmarks. Across six large real-world datasets, including PAMAP2, WISDM2, HARTH, ECG, SLEEP, and SKODA, Di-COT reaches 83.07% average linear-probe accuracy, versus 81.20% for CaTT and 80.69% for TNC, while training in 2.88 hours, faster than CaTT’s 3.47 hours and TNC’s 11.48 hours. On 1% labeled data it improves average accuracy to 76.36%, beating the best supervised baseline at 70.39% and TF-C at 73.55%. It also leads on 124 UCR and 28 UEA datasets with 81.33% and 71.24% average accuracy, respectively, and tops clustering with 0.508 NMI and 0.406 ARI. The key novelty is temporal contrast as sub-block classification: dense supervision, no augmentation bias, and multi-scale robustness from stochastic overlap.
Original abstract
Self-supervised learning for time-series representation aims to reduce reliance on labeled data while maintaining strong downstream performance, yet many existing approaches incur high computational costs or rely on assumptions that do not hold across diverse temporal dynamics. In this work, we introduce Divide and Contrast (Di-COT), an unsupervised framework that avoids data augmentation and multiple encoder passes by contrasting informative substructures within a window rather than individual timesteps. Di-COT stochastically partitions each window into a small number of overlapping sub-blocks per iteration, enabling efficient and meaningful contrast while mitigating false positives during temporal transitions. To further improve scalability, we adopt a contrastive objective whose computation depends on the batch size and the number of sub-blocks, making loss computation independent of sequence length. Extensive experiments on six large-scale real-world datasets, as well as the UCR and UEA benchmarks, demonstrate that Di-COT learns semantically structured and transferable representations, achieving state-of-the-art performance on classification, clustering, $k$NN, and cross-dataset transfer, while substantially reducing training time. The source code is publicly available at https://github.com/sfi-norwai/Di-COT.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.