Chiaroscuro Attention: Spending Compute in the Dark
AuthorsPrateek Kumar Sikdar
Resources
This paper asks when transformers really need full attention, and shows that a hybrid of spectral mixing and attention can cut compute while outperforming a strong baseline on some text tasks.
Key results
CHIAR DCT+Attn validation perplexity
Full-attention WikiText-103 baseline
Reduction for CHIAR DCT+Attn versus full attention
CHIAR DCT+Attn total GFLOPs
Full-attention GFLOPs on WikiText-103
What the paper found
Chiaroscuro Attention, or CHIAR-Former, proposes a 4-layer hybrid transformer that routes each token to DCT spectral mixing, RBF kernel mixing, or full self-attention using per-token spectral entropy as a complexity signal. The key finding is routing collapse: learned gates consistently reject RBF and converge to a DCT+attention design, which is not a failure but a discovery of the minimal useful operator set. On WikiText-103, that collapsed CHIAR DCT+Attn variant reaches 36.54 validation perplexity versus 66.62 for a full-attention baseline, a 45% improvement, while cutting attention FLOPs by 62.5% and total compute from 1.88 to 1.11 GFLOPs. The model’s operating regime is narrow: it excels on large, naturalistic text, matches a full-attention baseline on IMDB within 1.24 percentage points, but underperforms on WikiText-2 and ListOps, where scarce data or exact symbolic rule-following favor uniform attention. The paper’s central contribution is a theoretically grounded routing rule—spectral entropy derived from DCT energy concentration—that turns operator selection into a token-wise compute allocation problem rather than a uniform transformer design.
Original abstract
Standard transformers apply self-attention uniformly at every layer and token, regardless of whether the input requires dynamic cross-token interaction. We propose CHIAR-Former (Chiaroscuro Attention), a 4-layer hybrid transformer that routes each token to one of three operators - DCT spectral mixing, RBF kernel mixing, or full self-attention - based on per-token spectral entropy, a theoretically justified complexity signal. Through systematic ablation on WikiText-103, we discover routing collapse: the router consistently rejects RBF in favour of DCT and attention, revealing that spectral mixing and dynamic attention are complementary and sufficient. A purpose-designed DCT+Attention-only variant achieves Val PPL 36.54 on WikiText-103 - a 45% improvement over a full-attention baseline (PPL 66.62) at 62.5% fewer attention FLOPs. We extend evaluation to WikiText-2, IMDB sentiment classification, and synthetic ListOps operations, establishing a clear operating regime: CHIAR-Former excels on large-scale naturalistic text where token diversity supports spectral specialisation, while full attention retains an edge on small datasets and synthetic pattern-matching tasks. These findings - both the wins and the losses - together define when and why spectral routing earns its keep.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.