NTH

Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure

AuthorsShuyang Xiang

September 28, 2026 2 min read
Watch on YouTube
The one-line take

The study finds that real paragraph structure is revealed not merely by where attention concentrates, but by how deeply that compression develops across model layers.

Key results

-0.772
Code compression depth

Exact-distance residualized-attention depth U* for the Python Code corpus.

-0.481
WikiText-2 compression depth

Exact-distance residualized-attention depth U* for WikiText-2.

-0.305
OpenWebText compression depth

Exact-distance residualized-attention depth U* for OpenWebText.

1.522
Code fake-boundary response

Normalized fake-split response relative to the average genuine paragraph-boundary effect.

5000
Training steps

Training duration shared by the positional-encoding variants.

What the paper found

This study tests whether paragraph structure changes Transformer attention beyond ordinary token distance. It introduces hierarchical rotary positional encoding, or hRoPE, with independent paragraph, sentence, and token coordinates, then changes only the paragraph coordinate p1 while keeping tokens and reading order identical. Using 8-layer, 8-head models with dmodel 512, context length 1024, GPT-2 byte-pair encoding, and 5000 training steps, experiments cover WikiText-2, OpenWebText, and Python Code. An exact-distance residualized-attention estimator, retaining cells with at least 2000 token pairs and token distances up to 1024, finds compression near paragraph boundaries in all three corpora. However, a density-matched random p1 channel also causes compression, so the location of the minimum is not diagnostic; the signature is compression depth. The measured depths are −0.772 for Code, −0.481 for WikiText-2, and −0.305 for OpenWebText. Causal fake-merge and fake-split interventions produce opposite attention changes, while flat RoPE and sentence-only controls show no paragraph-coordinate effect. For Code, a fabricated boundary generated a normalized response of 1.522 times the average genuine-boundary effect. Across eight corpus-only measures spanning lexical persistence, paragraph length, and embedding coherence, none reproduces the cross-corpus depth ordering, although embedding coherence comes closest. The result supports hierarchical, coordinate-dependent attention rather than a complete geometric theory, and leaves the cause of corpus-dependent compression depth unresolved.

Original abstract

Standard positional encodings represent position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate p1, and measure cross-paragraph attention with a token-distance-exact estimator. Attention is compressed relative to a token-distance-matched baseline in every corpus, but compression alone is not diagnostic of true structure: an architecturally identical channel with density-matched random labels is compressed too, more shallowly. What distinguishes real structure is the depth of compression, which is greater and corpus-dependent while the control's is not. Comparing eight corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence), none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest. Compression depth, not its location, is the reproducible signature of genuine paragraph structure in our setting.

Read the original paper

More in Attention Mechanisms

Browse all 18 papers →
01Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis