NTH

Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors

AuthorsMantas Lukauskas

September 3, 2026 3 min read
Watch on YouTube
The one-line take

Prompt compression can save tokens in English but silently erase meaning in other languages, so multilingual users need much safer compression budgets.

Key results

10
Audit languages

The evaluation uses fully parallel Belebele items in 10 languages.

0.57
LLMLingua-2 English utilization

GPT-5.4-mini retains 0.57 normalized context utilization for English at a 0.33 keep-rate.

0.10
LLMLingua-2 Lithuanian utilization

GPT-5.4-mini retains only 0.10 normalized context utilization for Lithuanian at a 0.33 keep-rate.

-0.03
Chinese utilization

Chinese falls to -0.03 normalized utilization under LLMLingua-2 at a 0.33 keep-rate, below no-context utility.

92%
XProvence v2 empty Chinese contexts

XProvence v2 returns empty contexts for 92% of Chinese items at its aggressive threshold.

What the paper found

This controlled audit tests whether extractive prompt compression transfers beyond English, using 300 parallel Belebele reading-comprehension items in 10 languages, five scripts, four learned compressors, deterministic baselines, and 11 target models including OpenAI’s GPT-5.4-mini, Anthropic’s Claude Haiku 4.5, Google’s Gemini 3.5 Flash, Meta’s Llama 4 Maverick, DeepSeek V4 Flash, and Qwen 3.7 Plus. LLMLingua-2, trained with English supervision, appears safe at a 0.75 keep-rate but collapses at 0.33: English retains 0.57 normalized context utilization on GPT-5.4-mini, Lithuanian retains 0.10, and Chinese falls to -0.03, below the no-context baseline. The same transfer gap replicates across compressor backbones, including XLM-R, mBERT, and ModernBERT, and across target-model vendors, while budget-matched TF-IDF, truncation, random deletion, and lemmatization show no comparable cross-lingual cliff. This implicates supervision language rather than architecture. Multilingually supervised XProvence v1 eliminates the gap, but translated-data successor v2 deletes 92% of Chinese contexts at its aggressive threshold, exposing calibration and segmentation failures. In long-context retrieval-style tests, aggressive learned compression reduces several non-English contexts to no-context utility. A translate-then-compress pipeline using NLLB-200-distilled-600M reaches comparable or better quality in three of five languages at 0.18 relative token cost, versus 0.33 for native compression. The practical conclusion is that English-centric compressor scores overstate multilingual safety: use deterministic methods, multilingual calibration, or translation before compression, and monitor achieved rather than requested rates.

Original abstract

Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLingua-2 report strong results on English benchmarks. Most other languages already pay a token premium: the same content costs 1.3-1.8x more tokens than in English. We ask whether compression closes or widens this gap. Using fully parallel data in ten languages spanning five scripts, with controls budget-matched in the target model's tokenizer, we audit four learned compressors against four deterministic baselines, on eleven target models from ten vendors (over 250,000 evaluation calls). Three of the compressors are trained with English supervision (LLMLingua-2 XLM-R/mBERT; Kompress-v2 from the production Headroom stack); the fourth, XProvence, is trained multilingually. First, the transfer gap is real, replicates across target models and compressor backbones, and is strongly rate-dependent: at a 0.33 keep-rate English retains 57-62% of normalized context utilization while Lithuanian retains 10-24% and Chinese essentially none, despite Chinese having the smallest token premium. Second, the gap tracks compression supervision data, not architecture. All three English-trained compressors show it, deterministic methods show no comparable gap, and the multilingually trained XProvence v1 shows none. Its v2 release, retrained on translated data, empties 92% of Chinese contexts at its aggressive threshold without any warning. Third, in a harder long-context setting, aggressive learned compression drives compressed contexts to or below no-context utility in three of five non-English languages. A translate-then-compress pipeline matches or beats native compression at roughly half the token cost in three of five tested languages. We release all code, compressions, and model outputs. Safe compression budgets are much smaller outside English.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis