Site search
Search the briefings
Every paper we have covered, searchable by title, author, topic, or key result.
Latest briefs
- AttentionCoWindow Attention: Full Causal Coverage Is a Collective PropertyCoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.
- EfficiencyDecoding Looped Transformers Better for (Almost) FreeLoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
- LlmFinetuning with Sampling: SFT Learns Better Than You ThinkBy sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
- LlmGeneralization Dynamics of LM Pre-trainingLanguage models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
- AttentionHLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic MixingHLA makes linear attention more selective by letting each query dynamically choose which compressed chunks of long-context history to access.
- AttentionMinkowskiPE: Minkowski Positional Encoding for Spatiotemporal PerceptionMinkowskiPE gives attention a physics-inspired sense of spacetime, improving both molecular dynamics and video prediction with far fewer parameters.
Browse by topic
Large Language Models 81Generative Models 63Multimodal AI 61Computer Vision 58Diffusion Models 58AI Agents 56Efficient AI 55Reinforcement Learning 54Robotics 50Embodied AI 48Foundation Models 47AI Benchmarks 45AI for Science 43Code Generation 43Transformers 42World Models 41AI Reasoning 39AI Safety 39Optimization 36AI Hardware 34Graph Learning 32Speech AI 27Natural Language Processing 26Continual Learning 24Neural Networks 22Self-Supervised Learning 22Attention Mechanisms 18