Learning to Learn a Language
AuthorsLennart Carstens-Behrens, Holger Fröhlich
AffiliationsFraunhofer Institute for Algorithms and Scientific Computing SCAI · University Hospital Bonn
A transformer trained on synthetic worlds learns to infer the hidden rules of real language and other sequential data without ever seeing language during training.
Key results
Measured size of the byte-level transformer.
Total tokens sampled from the synthetic prior.
Upper end of the 0.9–2.4 bits-per-byte range at 1M bytes of context.
Bits per symbol after 1M digits of context.
PFLM bits per byte on source code, lower than the tested compressors.
What the paper found
The Prior-Fitted Language Model, or PFLM, tests whether a model can learn to infer a language rather than memorize one: its 295M-parameter byte-level transformer is trained only on synthetic sequences from freshly sampled recurrent structural causal models, then predicts real data from context with frozen weights. The prior is designed to reproduce natural text’s Zipfian frequencies and long-range dependencies, while each sequence comes from a different synthetic generator. After training on 150B synthetic tokens, PFLM lowers its prediction cost on Wikipedia in English, Chinese, Hindi, Arabic, Japanese, and Korean from 8 bits per byte to between 0.9 and 2.4 bits per byte at 1M bytes of context. It also learns to count, compare magnitudes, and estimate sums from numeral examples; on the Rudin–Shapiro sequence, its cost reaches 0.04 bits per symbol after 1M digits. Across six non-text domains, it beats gzip and PPMd; for source code, it achieves 1.54 bits per byte. The model uses a hybrid attention design similar to Qwen3-Next and was trained on NVIDIA H100 GPUs. In contrast with GPT-3’s text-based pretraining, these results suggest that carefully constructed statistical priors can support in-context learning without exposure to real language during training.
Original abstract
We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only way to predict the continuation is to infer the language from the prefix. Samples from this prior share the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence. On Wikipedia in six languages, bits per byte fall from the uniform eight to between 0.9 and 2.4 at one million bytes of context. Given numerals instead of text, PFLM learns to count, to compare magnitudes, and to add approximately. It predicts deterministic sequences like Rudin-Shapiro or the prime indicator, and it compresses six non-text domains, from source code to speech, below gzip and PPMd. The model has not learned a language. It has learned to learn one.
Read the original paperMore in Large Language Models
Browse all 84 papers →HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu, Jianquan Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, Haizhou Li, Benyou Wang
OnePO uses temporary teacher guidance and reinforcement learning alone to turn general LLMs into stronger medical specialists without the usual supervised fine-tuning stage.
On-Policy Distillation with Negative-Policy Rollouts
Jaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo
NP-OPD improves language-model distillation by exposing students to rollouts from a weaker policy, helping them learn from the teacher while moving away from inferior behavior.
Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.