Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure
AuthorsShuyang Xiang
Resources
The study finds that real paragraph structure is revealed not merely by where attention concentrates, but by how deeply that compression develops across model layers.
Key results
Exact-distance residualized-attention depth U* for the Python Code corpus.
Exact-distance residualized-attention depth U* for WikiText-2.
Exact-distance residualized-attention depth U* for OpenWebText.
Normalized fake-split response relative to the average genuine paragraph-boundary effect.
Training duration shared by the positional-encoding variants.
What the paper found
This study tests whether paragraph structure changes Transformer attention beyond ordinary token distance. It introduces hierarchical rotary positional encoding, or hRoPE, with independent paragraph, sentence, and token coordinates, then changes only the paragraph coordinate p1 while keeping tokens and reading order identical. Using 8-layer, 8-head models with dmodel 512, context length 1024, GPT-2 byte-pair encoding, and 5000 training steps, experiments cover WikiText-2, OpenWebText, and Python Code. An exact-distance residualized-attention estimator, retaining cells with at least 2000 token pairs and token distances up to 1024, finds compression near paragraph boundaries in all three corpora. However, a density-matched random p1 channel also causes compression, so the location of the minimum is not diagnostic; the signature is compression depth. The measured depths are −0.772 for Code, −0.481 for WikiText-2, and −0.305 for OpenWebText. Causal fake-merge and fake-split interventions produce opposite attention changes, while flat RoPE and sentence-only controls show no paragraph-coordinate effect. For Code, a fabricated boundary generated a normalized response of 1.522 times the average genuine-boundary effect. Across eight corpus-only measures spanning lexical persistence, paragraph length, and embedding coherence, none reproduces the cross-corpus depth ordering, although embedding coherence comes closest. The result supports hierarchical, coordinate-dependent attention rather than a complete geometric theory, and leaves the cause of corpus-dependent compression depth unresolved.
Original abstract
Standard positional encodings represent position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate p1, and measure cross-paragraph attention with a token-distance-exact estimator. Attention is compressed relative to a token-distance-matched baseline in every corpus, but compression alone is not diagnostic of true structure: an architecturally identical channel with density-matched random labels is compressed too, more shallowly. What distinguishes real structure is the depth of compression, which is greater and corpus-dependent while the control's is not. Comparing eight corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence), none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest. Compression depth, not its location, is the reproducible signature of genuine paragraph structure in our setting.
Read the original paperMore in Attention Mechanisms
Browse all 18 papers →CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
Zhuokun Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang
HLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
Yuhao Li, Louie Hong Yao, Tianyi Shi, Hanqun Cao, Hongxia Hao, Zhen Zhao, Shengchao Liu
MinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.