NTH

Chiaroscuro Attention: Spending Compute in the Dark

AuthorsPrateek Kumar Sikdar

June 17, 2026 2 min read
Watch on YouTube
The one-line take

This paper asks when transformers really need full attention, and shows that a hybrid of spectral mixing and attention can cut compute while outperforming a strong baseline on some text tasks.

Key results

36.54
WikiText-103 val PPL

CHIAR DCT+Attn validation perplexity

66.62
Baseline val PPL

Full-attention WikiText-103 baseline

62.5%
Attention FLOP reduction

Reduction for CHIAR DCT+Attn versus full attention

1.11
Total compute

CHIAR DCT+Attn total GFLOPs

1.88
Baseline total compute

Full-attention GFLOPs on WikiText-103

What the paper found

Chiaroscuro Attention, or CHIAR-Former, proposes a 4-layer hybrid transformer that routes each token to DCT spectral mixing, RBF kernel mixing, or full self-attention using per-token spectral entropy as a complexity signal. The key finding is routing collapse: learned gates consistently reject RBF and converge to a DCT+attention design, which is not a failure but a discovery of the minimal useful operator set. On WikiText-103, that collapsed CHIAR DCT+Attn variant reaches 36.54 validation perplexity versus 66.62 for a full-attention baseline, a 45% improvement, while cutting attention FLOPs by 62.5% and total compute from 1.88 to 1.11 GFLOPs. The model’s operating regime is narrow: it excels on large, naturalistic text, matches a full-attention baseline on IMDB within 1.24 percentage points, but underperforms on WikiText-2 and ListOps, where scarce data or exact symbolic rule-following favor uniform attention. The paper’s central contribution is a theoretically grounded routing rule—spectral entropy derived from DCT energy concentration—that turns operator selection into a token-wise compute allocation problem rather than a uniform transformer design.

Original abstract

Standard transformers apply self-attention uniformly at every layer and token, regardless of whether the input requires dynamic cross-token interaction. We propose CHIAR-Former (Chiaroscuro Attention), a 4-layer hybrid transformer that routes each token to one of three operators - DCT spectral mixing, RBF kernel mixing, or full self-attention - based on per-token spectral entropy, a theoretically justified complexity signal. Through systematic ablation on WikiText-103, we discover routing collapse: the router consistently rejects RBF in favour of DCT and attention, revealing that spectral mixing and dynamic attention are complementary and sufficient. A purpose-designed DCT+Attention-only variant achieves Val PPL 36.54 on WikiText-103 - a 45% improvement over a full-attention baseline (PPL 66.62) at 62.5% fewer attention FLOPs. We extend evaluation to WikiText-2, IMDB sentiment classification, and synthetic ListOps operations, establishing a clear operating regime: CHIAR-Former excels on large-scale naturalistic text where token diversity supports spectral specialisation, while full attention retains an edge on small datasets and synthetic pattern-matching tasks. These findings - both the wins and the losses - together define when and why spectral routing earns its keep.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis