NTH

Understanding Data Temporality Impact on Large Language Models Pre-training

AuthorsPilchen Hippolyte, Fabre Romain, Signe Talla Franck, Perez Patrick, Grave Edouard

June 4, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that the order of pre-training data matters: training large language models on time-ordered text can make them better at knowing which facts were true when, without hurting general language ability.

Key results

6B
Model size

The paper pretrains matched 6B-parameter Transformer decoder models to isolate the effect of temporal ordering.

2.5T
Training tokens

Both the shuffled baseline and sequential curriculum are trained on a total of 2.5T Common Crawl tokens.

7167
KairosQA size

The paper introduces KairosQA, a temporally grounded question-answering benchmark with 7,167 subject-relation pairs.

What the paper found

Understanding Data Temporality Impact on Large Language Models Pre-training, by Romain Fabre, Hippolyte Pilchen, Franck Signe Talla, Patrick Perez, and Edouard Grave at Kyutai, argues that standard shuffled pre-training obscures when facts were learned and produces a measurable “knowledge horizon” lag. The authors pretrain matched 6B-parameter Transformer decoders on 2.5T Common Crawl tokens in two regimes: a globally shuffled baseline and a strictly chronological curriculum spanning 2018–2025, using AdamW, a 4096-token context, and 128 H100 GPUs. To evaluate temporal grounding, they introduce KairosQA, a 7,167-pair benchmark built from Wikidata subject–relation–year triplets, filtered for temporal variation, popularity, and ambiguity, then tested in both cloze and generative settings; they also compare against TAQA. On OLMES, sequential training matches shuffled training on general language understanding, reaching comparable final scores while changing the learning trajectory. On KairosQA, however, chronological ordering consistently yields fresher factual knowledge: sequential checkpoints peak in the year immediately before their cutoff and outperform shuffled models on recent facts, reversing the recency decay seen in open-source baselines such as Llama 3.1-8B, Gemma 3, Qwen3, and Olmo3. The effect is strongest on 2023–2024 questions, where shuffled models often approach random accuracy, while the sequential model retains a clear margin. The paper also shows that simple replay or model soup only partially mitigates catastrophic forgetting, indicating that temporal ordering improves factual freshness but does not fully solve historical retention.

Original abstract

Large language models (LLMs) are typically trained on shuffled corpora, yielding models whose knowledge is frozen at train time and whose temporal grounding remains poorly understood. In this work, we study the impact of pre-training dynamics on the acquisition of time-sensitive factual knowledge, focusing specifically on data ordering. Our main contributions are twofold. First, we introduce a comprehensive benchmark of over 7,000 temporally grounded questions and an evaluation protocol that enables analysis of whether models correctly associate facts with their corresponding time periods. Second, we pretrain 6B-parameter models on temporally ordered Common Crawl snapshots and compare them against standard shuffled pre-training. Our results show that sequentially trained models match shuffled baselines on general language understanding and common knowledge while consistently exhibiting more up-to-date and temporally precise knowledge. Temporally ordered pre-training yields improved factual freshness, while shuffled pre-training peaks on older data, possibly due to increased factual repetition. These findings, along with the release of our code at https://github.com/kyutai-labs/kairos , checkpoints, and datasets at https://huggingface.co/collections/kyutai/kairos provide a foundation for future research on continual learning for LLMs.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis