LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.
This work shows, in theory, how transformers can adapt to data living on locally different geometric structures and still achieve statistically optimal in-context prediction.
This paper shows that letting a Transformer spend extra computation on latent thought steps can approach the quality of doubling its depth while using substantially fewer parameters.
Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina
This work argues that Transformers can blend two thoughts in one computation and proposes a way to separate them into coherent parallel continuations.
This work studies how Transformer layers change representations by separating updates that reinforce existing directions from those that redirect them, revealing implications for editing, compression, and training.
This paper tests whether transformers really need positional encodings to recognize token distances, revealing when different schemes help or hurt generalization to unseen delays.
RenderFormer-V2 uses specialized transformer attention to render complex scenes with diverse materials, environments, and volumetric effects without retraining for each scene.
Jie Chen, Xiangqian Yu, Yanchao Lian, Tan Lu, Run Yang, Zhengchun Shang, Xing Wang, Cheng Chen, Ke Hu, Qiang Li, Tianjiu Yin, Xiaobing Liu
ReST adapts scalable Transformers for industrial recommendation by reusing user-history computation across candidates while improving ranking quality under strict latency constraints.
TokenMatch uses curvature-aware mesh tokens and attention to rapidly match corresponding points across difficult 3D shapes, including partial and heavily deformed ones.
The paper shows that putting a model’s reasoning trace before a long document can dramatically improve its ability to solve long-context reasoning tasks.
This paper replaces conventional attention with explicit Self-and-Exchange relations, aiming to preserve language-model quality while improving efficiency.
Maglev teaches a lightweight recurrent Transformer to preserve long-range context in a compact memory, enabling faster long-context generation without full attention at inference.
Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
A transformer that feeds its hidden thoughts back into the next decoding step appears to gain better performance and shorter reasoning with little extra cost.
Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo
This study shows that unusual activation spikes in hybrid-attention language models follow predictable patterns tied to where full attention occurs and how activations are canceled.
UniR² turns recommendation recall and ranking into two coordinated tasks within one decoder-only Transformer, aiming to improve consistency and efficiency in production systems.
Henry Ndubuaku, Karen Mosoyan, Jakub Mroz, Noah Cylich, Satyajit Kumar, Parkirat Sandhu, Roman Shemet, Justin H Lee
A large controlled study finds that attention alone can nearly reproduce transformer performance, though feed-forward layers still provide an advantage for recalling knowledge stored in the weights.
This paper recasts Transformer behavior as continuous geometry and thermodynamics, offering an intriguing but currently insufficiently substantiated mathematical lens.
Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan
xHC expands Transformer memory through many residual streams while using sparse updates and optimized memory traffic to make large-scale LLM training more efficient.
This paper turns transformer attention heads into small executable Python programs, offering a new way to explain and partially replace opaque neural behavior with human-readable code.
This paper asks when transformers really need full attention, and shows that a hybrid of spectral mixing and attention can cut compute while outperforming a strong baseline on some text tasks.
Ziqing Qiao, Yinuo Xu, Chaojun Xiao, Zhou Su, Zihan Zhou, Yingfa Chen, Xiaoyue Xu, Xu Han, Zhiyuan Liu
This paper shows that in hybrid language models, efficient attention mostly shapes how quickly long-context skills emerge, while full attention does the heavy lifting for retrieval.
Zhaofeng Wu, Oliver Sieberling, Shawn Tan, Rameswar Panda, Yury Polyanskiy, Yoon Kim
This paper proposes making transformers wide at the edges and narrow in the middle, showing that smarter width allocation can improve language model efficiency and performance.
SISA is a new attention design that lets state-space models guide which tokens matter inside softmax attention, aiming to combine global retrieval with better prioritization in language models.
Jiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, Daniel Cremers
This paper turns attention from pairwise token matching into a functional map over continuous fields, aiming to make transformer-style operator learning more compact, resolution-invariant, and better suited for PDEs and other scientific tasks.
This paper makes transformer feedforward layers both smaller and more interpretable by replacing dense expert blocks with sparsely selected single-neuron linear experts.
This paper asks whether transformers really need separate query, key, and value projections—and finds that tying some of them can preserve much of the quality while greatly reducing memory use.
Parallax is a new attention design for language models that makes local linear attention practical at scale and claims better training quality and faster decoding than FlashAttention-style baselines.
Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski
This paper separates what Transformers know about meaning and word order into distinct streams, revealing how positional structure is stored and showing that the split can improve representations.
This paper shows that transformers can go beyond point estimates and learn full Bayesian predictive distributions in context, with theory explaining why architecture choices like normalization and attention depth matter.
Maxime Meyer, Mario Michelessa, Caroline Chaux, Vincent Y. F. Tan
This paper shows that transformers can only generate a surprisingly limited set of outputs, with the reachable sequence space growing linearly with prompt length and shrinking sharply beyond a threshold.
This paper shows that, in a simplified transformer model, large learning rates can make training settle into cycles or chaos instead of cleanly learning, revealing new stability thresholds for transformer dynamics.
Chunyuan Deng, Yizhe Zhang, Rui-Jie Zhu, Yuanyuan Xu, Jiarui Liu, T. S. Eugene Ng, Hanjie Chen
This paper makes looped transformers much cheaper by swapping in linear or sparse attention, showing you can keep quality while dramatically cutting compute.