NTH

Do Transformers Need Three Projections? Systematic Study of QKV Variants

AuthorsAli Kayyam, Anusha Madan Gopal, M Anthony Lewis

June 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper asks whether transformers really need separate query, key, and value projections—and finds that tying some of them can preserve much of the quality while greatly reducing memory use.

Key results

12
Tasks evaluated

The study systematically benchmarks projection-sharing variants across 12 tasks spanning synthetic, vision, and language modeling.

50%
300M KV cache reduction

For 300M parameter language models, Q-K=V reduces KV-cache memory by 50% relative to standard QKV attention.

3.1%
300M perplexity degradation

At 300M scale, Q-K=V incurs only a 3.1% perplexity degradation versus the QKV baseline.

2.48%
1.2B perplexity degradation

At 1.2B scale, Q-K=V shows a smaller 2.48% perplexity degradation versus QKV.

0.41%
Downstream accuracy drop

Across five 5-shot benchmarks at 1.2B scale, Q-K=V drops only 0.41 percentage points on average versus QKV.

What the paper found

This ICML 2026 paper, by Ali Kayyam, Anusha Madan Gopal, and M Anthony Lewis at BrainChip Inc., asks whether Transformers really need separate query, key, and value projections. The authors systematically compare three weight-tying variants: Q=K-V, Q-K=V, and Q=K=V, with optional 2D positional encoding to restore asymmetry in non-causal settings. Across 12 tasks spanning synthetic sequence manipulation, vision, and GPT-style language modeling on SlimPajama, they find that Q-K=V is the best projection-sharing design: in 300M-parameter language models it cuts KV-cache memory by 50% with only 3.1% perplexity degradation, and at 1.2B scale the gap shrinks to 2.48% while downstream 5-shot accuracy drops by just 0.41 percentage points on average across HellaSwag, PIQA, ARC-Easy, ARC-Challenge, and WinoGrande. By contrast, Q=K-V preserves training quality better than expected in vision tasks but gives no cache savings, while Q=K=V is too restrictive for language modeling, suffering a 25.4% perplexity hit at 300M scale. The study also shows projection sharing is orthogonal to head sharing: combining Q-K=V with GQA-4 yields 87.5% cache reduction, and with MQA it reaches 96.9%, enabling much lower-memory autoregressive inference. On 1.2B models, these architectural changes translate into measurable throughput gains on an A100, with Q-K=V improving decode throughput by 4.4–5.3% and reducing peak memory by 6.5–6.9% versus standard QKV.

Original abstract

Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value), b) Q=K-V (shared query-key), and c) Q=K=V (single projection). The last two variants produce symmetric attention maps; to address this, we also explore asymmetric attention via 2D positional encodings. Through experiments spanning synthetic tasks, vision (MNIST, CIFAR, TinyImageNet, anomaly), and language modeling (300M and 1.2B parameter models on 10B tokens), we discovered that our transformers perform on par or occasionally better than the QKV transformer. In language modeling, Q-K=V projection sharing achieves 50% KV cache reduction with only 3.1% perplexity degradation. Crucially, projection sharing is complementary to head sharing (GQA/MQA): combining Q-K=V with GQA-4 yields 87.5% cache reduction, while Q-K=V + MQA achieves 96.9%, enabling practical on-device inference. We show that Q-K=V preserves quality because keys and values can occupy similar representational spaces and attention operates in a low-rank regime, whereas Q=K-V breaks attention directionality. Our results systematically characterize projection sharing as an underexplored instance of weight tying in attention, with direct, quantifiable inference memory benefits, particularly valuable for edge deployment. The code is publicly available at https://github.com/Brainchip-Inc/Do-Transformers-Need-3-Projections

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis