NTH

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

AuthorsFanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel

August 22, 2026 2 min read
Watch on YouTube
The one-line take

LittleLearner is a language model trained only on elementary-school-level material, creating a controlled sandbox for studying how models gain knowledge and skills.

Key results

88B
LittleCurriculum size

Curated FineWeb-Edu tokens restricted to K–5 material.

5B
LittleLearner model size

Qwen3-based language model trained from scratch.

0%
Beyond-K–5 retention on CommonCoreText

Held-out evaluation found no retained Beyond-K–5 documents.

35%
K–5 retention on CommonCoreText

The precision-first filter preserved 35% of in-scope passages.

6.0%
Direct Beyond-K–5 accuracy

LittleLearner’s direct MathCAMPS accuracy.

5.9%
Natural-prose ICL Beyond-K–5 accuracy

Few-shot natural-prose demonstrations failed to improve out-of-scope reasoning.

What the paper found

LittleLearner introduces a controlled alternative to web-scale pretraining: LittleCurriculum, an 88B-token subset of FineWeb-Edu restricted to U.S. kindergarten-through-Grade-5 concepts, and LittleLearner, a 5B-parameter model trained from scratch with the Qwen3 architecture on NVIDIA B200 GPUs. Its precision-first pipeline combines Age-of-Acquisition filtering, FastText and ModernBERT classifiers, symbolic mathematics rules, and frequency sampling, with Gemini Flash judgments optimized through DSPy and OpenEvolve. On held-out CommonCoreText, the filter retains 0% of Beyond-K–5 material while preserving 35% of K–5 passages, establishing a sharp but deliberately conservative exposure boundary. The model remains competent on elementary language, science, and mathematics, but its performance declines sharply on advanced content; scaling from 0.6B to 5B parameters helps in-scope tasks and near-boundary problems, yet barely improves Grade-8 mathematics. Supervised fine-tuning followed by GRPO raises K–5 performance but does not restore Beyond-K–5 capabilities, and in-context learning shows the same limitation: on MathCAMPS, direct evaluation reaches 6.0% Beyond-K–5 accuracy, while natural-prose demonstrations reach 5.9%, despite improving K–5 accuracy from 34.0% to 36.8%. Compared with an unfiltered baseline and Gemma 2B, these results indicate that pretraining exposure, rather than scale, post-training, or prompting, sets the effective capability ceiling. The released corpus and model therefore function as a sandbox for studying genuine knowledge acquisition, continual learning, interpretability, calibration, and reinforcement-learning discovery beyond a known curriculum boundary.

Original abstract

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis