NTH
Research collection

Efficient AI research

Research on reducing training and inference costs through compression, pruning, quantization, and better computation. Compare savings alongside retained capabilities.

55 papers · Latest edition October 9, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Efficient AI papers

Newest editions first.

01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis
04Efficiency

Disaggregated Quantization: Specializing LLM Prefill and Decode

Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh

Disaggregated quantization gives LLM prefill and decode their own specialized weights and formats, improving low-bit accuracy while speeding up first-token generation.

Read analysis
06Efficiency

PoEM: Predicting RL Outcomes from Existing Policies

Kimia Hamidieh, Giannis Daras, Antonio Torralba

PoEM aims to mix the behaviors of already RL-trained models to predict how they would perform under a new reward, without running expensive RL again.

Read analysis
07Efficiency

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, :, Anyi Xu, B. Li, Bangcai Lin, Bing Xue, BingCheng Xian, Bingzheng Xu, Bochao Wu, Bowei Zhang, Boyi Deng, C. C. Yu, Chao Jin, Chaofan Lin, Chen Dong, Chenbing Wang, Chenfan Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyuan Zhang, Chenhao Xu, Chenqi Zhao, Chenze Shao, Chuhao Wang, Chuqi Zhang, Damai Dai, Dejian Yang, Deli Chen, Di Huang, Di Wu, Donghao Li, Erhang Li, Eric Fu, F. Zhou, Fangwei Zhou, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanglin Li, Guanting Chen, Guoai Cao, Guofan Fan, Guolai Meng, Guowei Li, Haichuan Zhang, Haiyang Ma, Haiyang Shen, Han Li, Han Yu, Han Zhang, Hangyuan Deng, Hanwei Xu, Hanxiang Xu, Hanxun Zhong, Hao Guo, Hao Jiang, Hao Li, Hao Qin, Haodong Wen, Haofen Liang, Haofeng Huang, Haohua Liu, Haoling Zhang, Haoming Luo, Haoran Yang, Haotian Xu, Haotian Yuan, Haoting Huang, Haowen Luo, Haoyang Cai, Haoyu Chen, Haozhe Ji, Hengran Zhang, Hengrui Wang, Hengxu Wu, Honghui Ding, Hongxuan Tang, Huadong Wang, Huanqi ...

DeepSeek-V4.1-Flash targets million-token agents by dramatically compressing KV caches while maintaining strong multimodal performance.

Read analysis
11Efficiency

Hyperparameter Scaling Laws Across MoE Sparsity

Changxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou

This work shows how to predict learning rates and batch sizes for increasingly sparse MoE models, making large-scale training more compute-efficient and transferable.

Read analysis
12Efficiency

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Heng Wang, Jielin Qiu, Wenting Zhao, Cheng Qian, Liangwei Yang, Jiawei Han, Heng Ji, Silvio Savarese, Shelby Heinecke, Huan Wang

Instead of carefully scoring which tokens to keep, this work shows that protecting the prompt and randomly evicting the rest can make long-form LLM reasoning substantially faster.

Read analysis
14Efficiency

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu

Qwen3.8-Next combines sparse experts, hybrid attention, external memory, and optimizer changes to achieve competitive performance with far less computation and improved training stability.

Read analysis
15Efficiency

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

Kairong Luo, Jiarui Cui, Yaorui Yin, Shengqi Chen, Yiming Yang, Linxiang Gao, Yanmohan Wang, Mingzhe Zhang, Kaiyue Wen, Kaifeng Lyu, Wenguang Chen

Puro-2B shows that a reasonably capable 2B language model can be pretrained from scratch for roughly $5,000 using consumer-grade GPUs and an open recipe.

Read analysis
16Efficiency

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Shuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu, Ion Stoica

FreeToken aims to make frontier-scale mixture-of-experts models practical on ordinary personal computers by dynamically coordinating CPU, GPU, memory, and model state.

Read analysis
19Efficiency

Matryoshka Language Model Suites

Nathan Godey, Yoav Artzi

One nested language model can serve as several differently sized models, cutting training costs while speeding up speculative decoding.

Read analysis
21Efficiency

Small-Scale Experiments: Are We There Yet?

Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi

Small models can predict large-model behavior after all—but only when their hyperparameters are tuned carefully enough to reveal the hidden scaling laws.

Read analysis
25Efficiency

Efficiency Matters in Autonomous Research

Haiqian Yang, Yuan Cao

This paper argues that autonomous research systems should be judged not only by what they discover, but also by how efficiently they find it.

Read analysis
27Efficiency

Token Reduction Is Not Cost Reduction

Sarel Weinberger, Amir Hozez

For coding agents, cutting tokens is not the same as cutting costs—and aggressive compression can even break successful software patches.

Read analysis
31Efficiency

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

Qingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Xiang Li, Ming Xu, Jiarui Li, Xiuyin Zhao

ShortOPD helps pruned language models regain useful free-form generation by training first on the prefixes they can handle and gradually extending their rollout length.

Read analysis
32Efficiency

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan Chen, Huanqi Cao, Wentao Zhang, Anyi Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang

DSpark speeds up LLM inference by generating draft tokens in a more structured way and verifying them adaptively, delivering major serving throughput gains in real production traffic.

Read analysis
37Efficiency

Linear Scaling Video VLMs for Long Video Understanding

Cristobal Eyzaguirre, Jiajun Wu, Juan Carlos Niebles

This paper makes long video understanding cheaper by replacing quadratic video attention with a stateful linear-time cache that preserves much of the accuracy of full self-attention.

Read analysis
39Efficiency

Efficient Training on Multiple Consumer GPUs with RoundPipe

Yibin Luo, Shiwei Gao, Huichuan Zheng, Youyou Lu, Jiwu Shu

RoundPipe is a new way to fine-tune huge language models efficiently on cheap GPUs by dynamically rotating work across devices to cut pipeline bubbles and boost throughput.

Read analysis
42Efficiency

Still: Amortized KV Cache Compaction in a Single Forward Pass

Charles O'Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara, Max Kirkby

Still makes long-context language models more practical by compressing their key-value cache in one fast forward pass, preserving performance even at extreme compression ratios.

Read analysis
43Efficiency

Can I Buy Your KV Cache?

Luoyuan Zhang

The paper argues that AI agents should stop recomputing the same prompt over and over, and instead buy and reuse precomputed KV caches to make LLM serving much cheaper.

Read analysis
44Efficiency

MiniMax Sparse Attention

Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao

MiniMax Sparse Attention makes ultra-long-context LLMs much faster by combining blockwise sparse retrieval with GPU-friendly kernels, delivering major inference speedups at 1M-token scale.

Read analysis
45Efficiency

End-to-End Context Compression at Scale

Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov

This paper introduces latent context models that compress long prompts into compact representations, making long-context language models faster and more memory-efficient without sacrificing as much quality.

Read analysis
46Efficiency

MobileMoE: Scaling On-Device Mixture of Experts

Yanbei Chen, Hanxian Huang, Ernie Chang, Jacob Szwejbka, Digant Desai, Zechun Liu, Vikas Chandra, Raghuraman Krishnamoorthi

MobileMoE shows that sparse mixture-of-experts language models can run efficiently on phones, delivering better speed-memory tradeoffs than dense on-device LLMs.

Read analysis
49Efficiency

Draft-OPD: On-Policy Distillation for Speculative Draft Models

Haodi Lei, Yafu Li, Haoran Zhang, Shunkai Zhang, Qianjia Cheng, Xiaoye Qu, Ganqu Cui, Bowen Zhou, Ning Ding, Yun Luo, Yu Cheng

Draft-OPD improves speculative decoding by training draft models on the kinds of errors they actually make at inference time, yielding faster LLM generation with better acceptance rates.

Read analysis
50Efficiency

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

Yiqun Liu, Yingsheng Wu, Ruqi Yang, Enrong Zheng, Honglei Qiu, Sijun He, Tai Liang, Jingjing Wu, Yuhan Zhou, Yiwei Zhang, Dongyan Chen, Weihan Yi, Xinqi Li, Siqi Bao

PassNet reframes LLM-based compiler optimization as graph pass generation, backing it with a large dataset and benchmark that show LLMs can meaningfully speed up real-world workloads.

Read analysis
51Efficiency

Approaching I/O-optimality for Approximate Attention

Pál András Papp, Aleksandros Sobczyk, Anastasios Zouzias

This paper makes attention much cheaper to run by reducing memory traffic, bringing it closer to the theoretical limits of efficiency.

Read analysis
54Efficiency

Pruning and Distilling Mixture-of-Experts into Dense Language Models

Junhyuck Kim, Jihun Yun, Haechan Kim, Gyeongman Kim, Joonghyun Bae, Jaewoong Cho

This paper shows how to turn large mixture-of-experts language models into smaller dense models that keep much of the performance while being easier and cheaper to deploy.

Read analysis