NTH

Learning to Learn a Language

AuthorsLennart Carstens-Behrens, Holger Fröhlich

AffiliationsFraunhofer Institute for Algorithms and Scientific Computing SCAI · University Hospital Bonn

October 10, 2026 2 min read
Watch on YouTube
The one-line take

A transformer trained on synthetic worlds learns to infer the hidden rules of real language and other sequential data without ever seeing language during training.

Key results

295M
PFLM parameters

Measured size of the byte-level transformer.

150B
Synthetic training tokens

Total tokens sampled from the synthetic prior.

2.4
Wikipedia prediction cost

Upper end of the 0.9–2.4 bits-per-byte range at 1M bytes of context.

0.04
Rudin–Shapiro prediction cost

Bits per symbol after 1M digits of context.

1.54
Source-code prediction cost

PFLM bits per byte on source code, lower than the tested compressors.

What the paper found

The Prior-Fitted Language Model, or PFLM, tests whether a model can learn to infer a language rather than memorize one: its 295M-parameter byte-level transformer is trained only on synthetic sequences from freshly sampled recurrent structural causal models, then predicts real data from context with frozen weights. The prior is designed to reproduce natural text’s Zipfian frequencies and long-range dependencies, while each sequence comes from a different synthetic generator. After training on 150B synthetic tokens, PFLM lowers its prediction cost on Wikipedia in English, Chinese, Hindi, Arabic, Japanese, and Korean from 8 bits per byte to between 0.9 and 2.4 bits per byte at 1M bytes of context. It also learns to count, compare magnitudes, and estimate sums from numeral examples; on the Rudin–Shapiro sequence, its cost reaches 0.04 bits per symbol after 1M digits. Across six non-text domains, it beats gzip and PPMd; for source code, it achieves 1.54 bits per byte. The model uses a hybrid attention design similar to Qwen3-Next and was trained on NVIDIA H100 GPUs. In contrast with GPT-3’s text-based pretraining, these results suggest that carefully constructed statistical priors can support in-context learning without exposure to real language during training.

Original abstract

We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only way to predict the continuation is to infer the language from the prefix. Samples from this prior share the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence. On Wikipedia in six languages, bits per byte fall from the uniform eight to between 0.9 and 2.4 at one million bytes of context. Given numerals instead of text, PFLM learns to count, to compare magnitudes, and to add approximately. It predicts deterministic sequences like Rudin-Shapiro or the prime indicator, and it compresses six non-text domains, from source code to speech, below gzip and PPMd. The model has not learned a language. It has learned to learn one.

Read the original paper

More in Large Language Models

Browse all 84 papers →
01Llm

HuatuoGPT-3: RL-Only Domain Adaptation from Base Models

Junying Chen, Xinyuan Xie, Ziniu Li, Wenyuan Gu, Jianquan Li, Xiang Wan, Guangjun Yu, Ruoyu Sun, Haizhou Li, Benyou Wang

OnePO uses temporary teacher guidance and reinforcement learning alone to turn general LLMs into stronger medical specialists without the usual supervised fine-tuning stage.

Read analysis
02Llm

On-Policy Distillation with Negative-Policy Rollouts

Jaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo

NP-OPD improves language-model distillation by exposing students to rollouts from a weaker policy, helping them learn from the teacher while moving away from inferior behavior.

Read analysis