Do Transformers Need Three Projections? Systematic Study of QKV Variants
AuthorsAli Kayyam, Anusha Madan Gopal, M Anthony Lewis
This paper asks whether transformers really need separate query, key, and value projections—and finds that tying some of them can preserve much of the quality while greatly reducing memory use.
Key results
The study systematically benchmarks projection-sharing variants across 12 tasks spanning synthetic, vision, and language modeling.
For 300M parameter language models, Q-K=V reduces KV-cache memory by 50% relative to standard QKV attention.
At 300M scale, Q-K=V incurs only a 3.1% perplexity degradation versus the QKV baseline.
At 1.2B scale, Q-K=V shows a smaller 2.48% perplexity degradation versus QKV.
Across five 5-shot benchmarks at 1.2B scale, Q-K=V drops only 0.41 percentage points on average versus QKV.
What the paper found
This ICML 2026 paper, by Ali Kayyam, Anusha Madan Gopal, and M Anthony Lewis at BrainChip Inc., asks whether Transformers really need separate query, key, and value projections. The authors systematically compare three weight-tying variants: Q=K-V, Q-K=V, and Q=K=V, with optional 2D positional encoding to restore asymmetry in non-causal settings. Across 12 tasks spanning synthetic sequence manipulation, vision, and GPT-style language modeling on SlimPajama, they find that Q-K=V is the best projection-sharing design: in 300M-parameter language models it cuts KV-cache memory by 50% with only 3.1% perplexity degradation, and at 1.2B scale the gap shrinks to 2.48% while downstream 5-shot accuracy drops by just 0.41 percentage points on average across HellaSwag, PIQA, ARC-Easy, ARC-Challenge, and WinoGrande. By contrast, Q=K-V preserves training quality better than expected in vision tasks but gives no cache savings, while Q=K=V is too restrictive for language modeling, suffering a 25.4% perplexity hit at 300M scale. The study also shows projection sharing is orthogonal to head sharing: combining Q-K=V with GQA-4 yields 87.5% cache reduction, and with MQA it reaches 96.9%, enabling much lower-memory autoregressive inference. On 1.2B models, these architectural changes translate into measurable throughput gains on an A100, with Q-K=V improving decode throughput by 4.4–5.3% and reducing peak memory by 6.5–6.9% versus standard QKV.
Original abstract
Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value), b) Q=K-V (shared query-key), and c) Q=K=V (single projection). The last two variants produce symmetric attention maps; to address this, we also explore asymmetric attention via 2D positional encodings. Through experiments spanning synthetic tasks, vision (MNIST, CIFAR, TinyImageNet, anomaly), and language modeling (300M and 1.2B parameter models on 10B tokens), we discovered that our transformers perform on par or occasionally better than the QKV transformer. In language modeling, Q-K=V projection sharing achieves 50% KV cache reduction with only 3.1% perplexity degradation. Crucially, projection sharing is complementary to head sharing (GQA/MQA): combining Q-K=V with GQA-4 yields 87.5% cache reduction, while Q-K=V + MQA achieves 96.9%, enabling practical on-device inference. We show that Q-K=V preserves quality because keys and values can occupy similar representational spaces and attention operates in a low-rank regime, whereas Q=K-V breaks attention directionality. Our results systematically characterize projection sharing as an underexplored instance of weight tying in attention, with direct, quantifiable inference memory benefits, particularly valuable for edge deployment. The code is publicly available at https://github.com/Brainchip-Inc/Do-Transformers-Need-3-Projections
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.