NTH
AI research

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization

AuthorsZhongzhu Zhou, Donglin Zhuang, Jisen Li, Ziyan Chen, Shuaiwen Leon Song, Ben Athiwaratkun, Xiaoxia Wu

May 19, 2026 3 min read
Watch on YouTube
The one-line take

OSCAR is a new way to shrink LLM KV caches to 2 bits by learning a better rotation from attention statistics, making long-context inference much faster and more memory-efficient without collapsing accuracy.

Key results

3.78 points
BF16 accuracy gap on Qwen3-4B-Thinking-2507

On the 32K-generation benchmark suite, OSCAR narrows the average accuracy gap to BF16 on Qwen3-4B-Thinking-2507.

1.42 points
BF16 accuracy gap on Qwen3-8B

On the same 32K-generation benchmark suite, OSCAR stays very close to BF16 on Qwen3-8B.

2.28 bits per KV element
OSCAR quantization budget

OSCAR’s main results are reported at an effective 2.28 BPE, including the BF16 sink/recent protection tokens.

128K tokens
Long-context robustness

OSCAR remains robust on RULER-NIAH up to 128K tokens, while rotation-only INT2 baselines collapse.

approximately 8×
KV-cache memory reduction

System measurements show OSCAR reduces KV-cache memory by about 8× relative to BF16.

up to 7×
Throughput speedup

At large batch sizes under the same memory budget, OSCAR achieves up to 7× higher throughput than BF16.

What the paper found

OSCAR is a 2-bit KV-cache quantization method that replaces data-oblivious rotations with offline, attention-aware spectral calibration. Rather than optimizing raw K/V reconstruction error, it estimates layer- and head-specific covariance targets from a small calibration pass: query covariance Q⊤Q for keys and score-weighted value covariance V⊤S⊤SV for values, where S is the attention matrix. It then builds fixed rotations RK = UQHHadPbr and RV = USHHadPbr from these eigenspaces, combining PCA, Walsh-Hadamard mixing, and bit-reversal permutation to equalize importance across channels and reduce outliers before asymmetric INT2 quantization with per-token percentile clipping. A key theoretical result shows these covariance targets minimize a frozen-error surrogate under diagonal residual assumptions. In production, OSCAR is implemented in SGLang with fused Triton kernels, paged KV-cache compatibility, BF16 protection for sink and recent tokens, and packed 2-bit history storage, so it fits standard long-context serving pipelines. Empirically, on 32K-generation evaluations over AIME25, GPQA-Diamond, HumanEval, LiveCodeBench v6, and MATH500, OSCAR at 2.28 bits per KV element cuts the BF16 accuracy gap to 3.78 points on Qwen3-4B-Thinking-2507 and 1.42 points on Qwen3-8B, while naive INT2 and QuaRot-INT2 collapse to near zero on these reasoning workloads; on Qwen3-32B and GLM-4.7-FP8, it is essentially on par with BF16. On RULER-NIAH up to 128K tokens, OSCAR stays robust where rotation-only INT2 fails, and system measurements show about 8× lower KV memory, up to 7× higher throughput at large batch sizes, and up to 3× faster batch-size-1 decoding versus BF16.

Original abstract

INT2 KV-cache quantization is attractive for long-context LLM serving, but it remains difficult to make both accurate and deployable. Simple rotations such as Hadamard transforms reduce outliers, but still degrade at INT2 because they are not aligned with downstream attention. We propose OSCAR, an Ultra-low-bit KV Cache quantization method that estimates attention-aware covariance structures offline and uses them to derive fixed rotations and clipping thresholds for quantization. In this way, it aligns KV quantization with the covariance structures that attention actually consumes. More importantly, we not only provide theoretical justification but also develop a fully deployable OSCAR system with a custom INT2 attention kernel that remains compatible with paged KV-cache serving and fused kernel pipelines, enabling seamless integration into modern LLM serving frameworks such as SGLang and vLLM. We evaluate our methods on recent reasoning models with reasoning traces of up to 32k tokens across 5 tasks. On Qwen3-4B-Thinking-2507 and Qwen3-8B, OSCAR reduces the BF16 accuracy gap to 3.78 and 1.42 points, respectively, while naive rotation INT2 collapses to nearly zero. We further scale OSCAR to Qwen3-32B and GLM-4.7 (358B params), where it remains effectively on par with BF16. On long context - RULER-NIAH up to 128K, OSCAR remains robust on both Qwen3 models, while naive rotation INT2 collapses. System-wise, OSCAR reduces KV-cache memory by approximately 8x, improves throughput by up to 7x at large batch sizes under the same memory budget, and accelerates batch-size-1 decoding by up to 3x over BF16 due to reduced memory bandwidth overhead.

Read the original paper