NTH

NITP: Next Implicit Token Prediction for LLM Pre-training

AuthorsXiangdong Zhang, Debing Zhang, Shaofeng Zhang, Xiaohan Qin, Yu Cheng, Junchi Yan

June 5, 2026 2 min read
Watch on YouTube
The one-line take

NITP tweaks LLM pretraining by making models predict not just the next token, but its hidden semantic representation too, improving downstream performance with little extra cost.

Key results

5.71
9B MoE MMLU-Pro gain

On the 9B MoE model, MMLU-Pro improves from 15.29 to 21.00 with NITP.

6.36
9B MoE C3 gain

On the 9B MoE model, C3 improves from 56.65 to 63.01 with NITP.

4.26
9B MoE CommonsenseQA gain

On the 9B MoE model, CommonsenseQA improves from 45.70 to 49.96 with NITP.

39.24
MTEB overall score

Frozen last-hidden-state quality on MTEB rises from 39.24 to 41.56 across 25 tasks.

2
Additional training FLOPs

The method adds approximately 2% additional training FLOPs.

0.5B
Model scale range

The method is evaluated across dense and MoE models ranging from 0.5B to 9B parameters.

What the paper found

NITP, or Next Implicit Token Prediction, is a new pre-training objective for large language models that augments standard next-token prediction with dense supervision in hidden-state space. The paper shows that ordinary next-token prediction provides sparse, one-hot gradients that leave most latent degrees of freedom under-constrained, causing last-layer representations to become anisotropic and low-rank during training. NITP fixes this by asking the model to predict an “implicit token,” defined as the stop-gradient shallow-layer representation of the next token, and aligning the final hidden state to it with a cosine-similarity loss plus the usual cross-entropy. The key design choice is temporal shift: predicting the next token’s semantic representation works, while same-position alignment collapses and hurts performance. Theoretical analysis shows that NITP adds positive curvature to the angular subspace of the loss Hessian, mitigating the semantic null space left by next-token prediction. Empirically, the method scales across Dense and MoE models from 0.5B to 9B parameters, including DeepSeek-V2-style MoE backbones, with roughly 2% extra training FLOPs and zero inference overhead. On the 9B MoE model, NITP improves MMLU-Pro by 5.71 points, C3 by 6.36 points, and CommonsenseQA by 4.26 points, while also raising frozen-representation quality on MTEB from 39.24 to 41.56 across 25 tasks. The authors position NITP as a compute-efficient way to make pre-training shape not just token accuracy, but the geometry and transferability of internal representations.

Original abstract

Standard next-token prediction (NTP) supervises language models solely through discrete labels in the output logit space. We argue that this sparse one-hot supervision leaves the latent representation space under-constrained, allowing hidden states to drift into degenerate and anisotropic configurations that can limit generalization. To address this issue, we propose Next Implicit Token Prediction (NITP), which augments discrete prediction with dense continuous supervision directly in the representation space. NITP trains the model to predict the implicit semantic content of the next token, using shallow-layer representations from the same model as stable self-supervised targets. We provide theoretical analysis showing that NITP regularizes the optimization landscape by mitigating under-constrained degrees of freedom and encouraging a compact, structured representation geometry. Empirically, across dense and MoE models ranging from 0.5B to 9B parameters, NITP consistently improves downstream performance with negligible computational overhead. On a 9B MoE model, NITP achieves a 5.7% absolute improvement on MMLU-Pro, along with gains of 6.4% on C3 and 4.3% on CommonsenseQA, with approximately 2% additional training FLOPs and no additional inference cost. Our implementation is available at https://github.com/aHapBean/NITP.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis