LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
AuthorsFanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
Resources
LittleLearner is a language model trained only on elementary-school-level material, creating a controlled sandbox for studying how models gain knowledge and skills.
Key results
Curated FineWeb-Edu tokens restricted to K–5 material.
Qwen3-based language model trained from scratch.
Held-out evaluation found no retained Beyond-K–5 documents.
The precision-first filter preserved 35% of in-scope passages.
LittleLearner’s direct MathCAMPS accuracy.
Few-shot natural-prose demonstrations failed to improve out-of-scope reasoning.
What the paper found
LittleLearner introduces a controlled alternative to web-scale pretraining: LittleCurriculum, an 88B-token subset of FineWeb-Edu restricted to U.S. kindergarten-through-Grade-5 concepts, and LittleLearner, a 5B-parameter model trained from scratch with the Qwen3 architecture on NVIDIA B200 GPUs. Its precision-first pipeline combines Age-of-Acquisition filtering, FastText and ModernBERT classifiers, symbolic mathematics rules, and frequency sampling, with Gemini Flash judgments optimized through DSPy and OpenEvolve. On held-out CommonCoreText, the filter retains 0% of Beyond-K–5 material while preserving 35% of K–5 passages, establishing a sharp but deliberately conservative exposure boundary. The model remains competent on elementary language, science, and mathematics, but its performance declines sharply on advanced content; scaling from 0.6B to 5B parameters helps in-scope tasks and near-boundary problems, yet barely improves Grade-8 mathematics. Supervised fine-tuning followed by GRPO raises K–5 performance but does not restore Beyond-K–5 capabilities, and in-context learning shows the same limitation: on MathCAMPS, direct evaluation reaches 6.0% Beyond-K–5 accuracy, while natural-prose demonstrations reach 5.9%, despite improving K–5 accuracy from 34.0% to 36.8%. Compared with an unfiltered baseline and Gemma 2B, these results indicate that pretraining exposure, rather than scale, post-training, or prompting, sets the effective capability ceiling. The released corpus and model therefore function as a sandbox for studying genuine knowledge acquisition, continual learning, interpretability, calibration, and reinforcement-learning discovery beyond a known curriculum boundary.
Original abstract
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.