NTH
Research collection

Large Language Models research

Explore how language models are trained, evaluated, and adapted for useful tasks. Compare reported capabilities, training choices, and limitations.

81 papers · Latest edition October 9, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Large Language Models papers

Newest editions first.

02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis
06Llm

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li

Latent-MOPD teaches one language model to combine several specialists by learning both what they predict and how they represent it.

Read analysis
07Llm

Periscope: Extending Frozen Language Models Beyond Their Context Window

Mohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi, Abdulrahman Mousa, Abdallah Ahmed, Tanveer Hussain, Naeemullah Khan

Periscope lets frozen language models search million-token documents on a single GPU by replacing one enormous read with a smart grid of small probes.

Read analysis
08Llm

Smaller Models, Better Rejects: Preference Distillation Scaling

Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov, Kayhan Behdin, Zhengze Zhou, Ran He, Alborz Geramifard, Rohit Jain, Zhe Zhao

Smaller models can generate more useful negative examples than students themselves, making preference training for large language models cheaper and potentially better.

Read analysis
10Llm

A Zeroth-Order Paradigm for LLM Preference Alignment

Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin

ComPO aligns language models using preference comparisons rather than direct likelihood optimization, aiming to reduce failures caused by small preference margins.

Read analysis
12Llm

Where Should a Document Live: Context, Representations, or Parameters?

Nathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert

This study asks whether new knowledge belongs in an LLM’s context, hidden representations, or weights, finding that KV-cache-based Cartridges often deliver the best accuracy but can cause forgetting.

Read analysis
13Llm

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih

A new distillation strategy uses teacher confidence to help language models gain reasoning skills without sacrificing the factual knowledge they learn during mid-training.

Read analysis
14Llm

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Zixuan Fu, Bingxiang He, Yuxin Zuo, Haohuan Huang, Jinqian Zhang, Ruhang Xiao, Cheng Qian, Qinyu Luo, Huan-ang Gao, Yudong Wang, Zhiyuan Liu, Ning Ding, Chaojun Xiao

This study finds that on-policy distillation of language models may need far fewer examples than expected, but many more training steps to fully absorb the supervision.

Read analysis
15Llm

StudentSim: Training LLM-based Student Simulators

Ke Yang, Chenglong Wang, Michel Galley, Chandan Singh, Jeevana Priya Inala, ChengXiang Zhai, Jianfeng Gao

StudentSim teaches LLMs to imitate individual learners and respond realistically to tutoring, enabling more personalized AI tutors.

Read analysis
16Llm

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Katherine Van Koevering, Anjalie Field

The study finds that everyday linguistic styles associated with women can cause LLMs to produce shorter and less formal workplace writing, revealing a subtle but consequential form of bias.

Read analysis
21Llm

MetaStrategy: Generative Ranking with Executable LLM Strategies

Chengyu Lai, Jiuning Lin, Zhibo Xiao, Xiaodong Zhu, Ruiquan Lan, Bin Zhang, Zihong Huang, Wendong Zhang, Chuxin Chen, Yinjiang Cai, Shuai Zhong, Lingqing Zhang, Dimin Wang, Jialin Zhu, Han Zhu

MetaStrategy uses a compact LLM to generate safe, executable ranking policies that improve real-world recommendation outcomes without adding online latency.

Read analysis
23Llm

Memory for Large Language Models

Sining Zhoubian, Dan Zhang, Evgeny Kharlamov, Jie Tang

This survey explains how modern language models remember, update, and retrieve information, and organizes the field’s many competing memory designs into one coherent framework.

Read analysis
24Llm

Why Large Language Models Fail at Tabular Prediction

Marta Garnelo, Wojciech M. Czarnecki

LLMs can reason impressively over many modalities, but this study finds that their tabular prediction ability collapses as the number of input dimensions grows.

Read analysis
25Llm

Co-Evolving LLM Evaluators and Policies via DynamicRubric

Beining Wang, Weihang Su, Hongtao Tian, Hao Kong, Tao Yang, Ting Yao, Qingyi Pan, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu

DynamicRubric keeps LLM evaluators useful as policies improve by generating adaptive rubrics tailored to each set of competing responses.

Read analysis
27Llm

Zing: Social Mind for LLMs

Zing Team, Ao Xiang, Bi Jingping, Chen Jiahui, Chen Lehan, Chen Yilin, Cheng Xueqi, Fan Yixing, Gan Kairong, Gao Haowen, Gao Jinhua, Gao Shuxuan, Gong Chang, Guo Jiafeng, Guo Ruijie, Han Zhouyu, He Guangfu, He Yichun, Jiang Shuo, Jing Shaoling, Jing Ya, Lei Chenhao, Lei Yan, Li Anqi, Li Chengao, Li Haoyu, Li Shitian, Liang Xinjian, Liu Zhaoge, Lyu Xingyu, Nie Zhuwei, Pang Liang, Quan Zeping, Shan Shiguang, Shen Huawei, Tang Xinran, Tian Feng, Wang Qian, Wang Ruiping, Wang Xiaohong, Xia Zaiyu, Xiao Yi, Xu Jiayuan, Xu Kehan, Xu Qianqian, Xu Tianyu, Xu Yongjun, Yang Haoming, Yang Jun, Yao Di, Yu Xiaoming, Zhang Futong, Zhang Jie, Zhang Shixuan, Zhang Yuxuan, Zhao Xinyu, Zhao Zhuoran, Zhong Yunfei, Zhu Shengyu

Zing gives LLMs a way to measure, learn, and apply social intelligence through specialized benchmarks, training, and runtime support.

Read analysis
29Llm

Distilled Reinforcement Learning for LLM Post-training

Chen Wang, Zhaochun Li, Jionghao Bai, Yining Zhang, Hexuan Deng, Ge Lan, Yue Wang

Distilled RL combines reinforcement learning with selective teacher guidance to help LLMs acquire new knowledge more effectively than standard RL or direct logit matching.

Read analysis
31Llm

Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?

Nishant Aggarwal, Ayushi Dubal, Sreeraj Kannakarankodi, Ian McDougall, Adarsh Mittal, Vishnu Ramadas, Noah Scott, Ranganath Selagamsetty, Weichu Yang, Karthikeyan Sankaralingam

This study finds that a carefully structured team of LLM reviewers can often critique computer architecture papers more deeply than individual human analysts, while still struggling with trust and calibration.

Read analysis
33Llm

Spectral Rewiring for Exploration, Purification, and Model Merging

Zhilong Zhang, Hongli Yu, Huan-ang Gao, Hanlin Wu, Yuxuan Song, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou

SAR trims language-model updates down to their most useful spectral directions, aiming to preserve reasoning while improving exploration, capability consolidation, and model merging.

Read analysis
35Llm

Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation

Xingyu Su, Jacob Helwig, Shubham Parashar, Atharv Chagi, Lakshmi Jotsna, Degui Zhi, James Caverlee, Dileep Kalathil, Shuiwang Ji

This paper shows how to turn an autoregressive language model into a diffusion-style language model much more efficiently by training it on its own generated trajectories while distilling knowledge from the original model.

Read analysis
39Llm

On the Geometry of On-Policy Distillation

Zhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong, Zhilin Wang, Yanxu Chen, Rongduo Han, Sunbowen Lee, Yi R. Fung

This paper shows that on-policy distillation for language models follows its own unique path in parameter space, with updates quickly locking into a low-dimensional subspace that still preserves performance.

Read analysis
43Llm

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

Xiang Hu, Xinyu Wei, Hao Gu, Minshen Zhang, Tian Liang, Huayang Li, Lei Zhu, Yan Wang, Sirui Han, Yushi Bai, Kewei Tu, Haitao Mi, Leo Liang

HiLS Attention is a new way to make LLMs handle much longer contexts by learning which chunks to attend to, cutting compute while keeping performance strong.

Read analysis
44Llm

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, Jinhao Dong, Zhifang Sui, Fuli Luo

MOPD is a new way to merge multiple specialized LLM skills by distilling several RL-trained teachers into one student using the student’s own rollouts, and it appears to work well at frontier scale.

Read analysis
45Llm

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

Yonggan Fu, Lexington Whalen, Abhinav Garg, Chengyue Wu, Maksim Khadkevich, Nicolai Oswald, Enze Xie, Daniel Egert, Sharath Turuvekere Sreenivas, Shizhe Diao, Chenhan Yu, Ye Yu, Weijia Chen, Sajad Norouzi, Jingyu Liu, Shiyi Lan, Ligeng Zhu, Jin Wang, Jindong Jiang, Morteza Mardani, Mehran Maghoumi, Song Han, Ante Jukić, Nima Tajbakhsh, Jan Kautz, Pavlo Molchanov

This paper introduces a language model that can switch between autoregressive, diffusion, and self-speculative decoding to trade off speed and quality, achieving major throughput gains over current open-source models.

Read analysis
46Llm

CausalMix: Data Mixture as Causal Inference for Language Model Training

Zinan Tang, Yukun Zhang, Shaomian Zheng, Zhuoshi Pan, Qizhi Pei, Dingnan Jin, Jun Zhou, Yujun Wang, Biqing Huang

CausalMix treats LLM data mixing like a causal problem, using learned treatment effects to choose better training mixtures that scale to larger models and data pools.

Read analysis
48Llm

Tapered Language Models

Reza Bayat, Ali Behrouz, Aaron Courville

This paper argues that language models should spend more capacity in earlier layers and less in later ones, and shows that a tapered width schedule can improve performance without increasing cost.

Read analysis
49Llm

AsyncOPD: How Stale Can On-Policy Distillation Be?

Wonjun Kang, Kevin Galim, Seunghyuk Oh, Minjun Kang, Sanghyun Park, Donghoon Kim, Minjae Lee, Minseo Kim, Rishabh Tiwari, Yuchen Zeng, Hyung Il Koo, Kangwook Lee

AsyncOPD shows how to train language models faster with asynchronous on-policy distillation without losing accuracy, by carefully handling stale rollouts and teacher-feedback estimates.

Read analysis
50Llm

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

Jian Yang, Shawn Guo, Wei Zhang, Tianyu Zheng, Yaxin Du, Haau-Sing Li, Jiajun Wu, Yue Song, Yan Xing, Qingsong Cai, Zelong Huang, Chuan Hao, Ran Tao, Xianglong Liu, Wayne Xin Zhao, Mingjie Tang, Weifeng Lv, Ming Zhou, Bryan Dai

LoopCoder-v2 shows that, for parallel loop transformers, two loops are the sweet spot: enough extra computation to boost code and agent performance, but not so much that positional mismatch overwhelms the gains.

Read analysis
51Llm

Predictable GRPO: A Closed-Form Model of Training Dynamics

Rajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta

This paper turns GRPO training into a predictable dynamical system, giving a closed-form explanation for reward growth, stability, and failure modes in LLM reasoning training.

Read analysis
53Llm

Improved Large Language Diffusion Models

Shen Nie, Qiyang Min, Shaoxuan Xu, Zihao Huang, Yuxuan Song, Yong Shan, Yankai Lin, Wayne Xin Zhao, Chongxuan Li, Ji-Rong Wen

This work shows that large language models can be trained with masked diffusion instead of next-token prediction and still become highly competitive on reasoning, math, and coding tasks.

Read analysis
54Llm

When is Your LLM Steerable?

Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou

This paper shows that you can predict whether an LLM steering trick will work by looking at its early hidden states, then use that prediction to search for better steering settings much more cheaply.

Read analysis
55Llm

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

Xuanming Zhang, Sining Zhoubian, Yuxuan Chen, Tianyi Tang, An Yang, Sean Du, Chujie Zheng, Fei Huang, Dayiheng Liu, Gao Huang, Jingren Zhou

This paper shows that for some aligned LLMs, the best next-token predictions may come from a near-final layer rather than the last one, and it uses that insight to improve reasoning with almost no extra cost.

Read analysis
57Llm

Superficial Beliefs in LLM Decision-Making

Gabriel Freedman, Francesca Toni

This paper shows that LLMs can make structured choices for hidden reasons that their own explanations only partly reveal.

Read analysis
58Llm

Re-Centering Humans in LLM Personalization

Lechen Zhang, Jiarui Liu, Tal August

This paper shows that LLM personalization looks much better on synthetic tests than on real human conversations, and introduces human-grounded data and evaluation methods to expose the gap.

Read analysis
59Llm

REVES: REvision and VErification--Augmented Training for Test-Time Scaling

Yuanxin Liu, Ruida Zhou, Xinyan Zhao, Amr Sharaf, Hongzhou Lin, Arijit Biswas, Mohammad Ghavamzadeh, Zhaoran Wang, Mingyi Hong

REVES improves LLM test-time reasoning by learning from near-miss answers, teaching models not just to solve problems but to recover from mistakes more effectively.

Read analysis
61Llm

Large Language Models Are Overconfident in Their Own Responses

Mario Sanz-Guerrero, Manuel Mager, Katharina von der Wense

This paper shows that LLMs are strangely more confident in answers they think are their own, and that a simple prompt trick can make them better calibrated without retraining.

Read analysis
62Llm

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

Yuhang Zhou, Lizhu Zhang, Yifan Wu, Mingyi Wang, Peng Bo, Jiayi Liu, Xiangjun Fan, Zhuokai Zhao

OmniOPD is a new way to train smaller models from stronger teachers without needing token logits, using chunk-level semantic verification and uncertainty-based checks to improve distillation, especially on math tasks.

Read analysis
64Llm

Less is More: Early Stopping Rollout for On-Policy Distillation

Zhou Ziheng, Jiaqi Li, Huacong Tang, Ying Nian Wu, Demetri Terzopoulos

This work shows that for on-policy distillation, letting the student only roll out the first few response tokens can actually outperform full rollouts, while being faster and more stable to train.

Read analysis
66Llm

Trajectory-Refined Distillation

Li Jiang, Haoran Xu, Yichuan Ding, Amy Zhang

This paper improves LLM distillation by fixing bad reasoning prefixes at the trajectory level, helping smaller models learn better from teacher-guided rollouts.

Read analysis
67Llm

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

Tao Chen, Gangwei Jiang, Pengyu Cheng, Siyuan Huang, Yihao Liu, Jingwei Ni, Jiaqi Guo, Mengyu Zhou, Kai Tang, Junling Liu, Qinliang Su, Xiaoxi Jiang, Guanjun Jiang

Skill-RM turns reward evaluation into an agent-like skill that dynamically combines rules, references, and rubrics to judge LLM outputs more consistently.

Read analysis
77Llm

Rethinking the Multilingual Reasoning Gap with Layer Swap

Maxence Lasbordes, Amélie Chatelain, Djamé Seddah

This paper shows that multilingual reasoning in LLMs can be improved by swapping in stronger mid-layers from an English model, revealing a shared reasoning core beneath language-specific outer layers.

Read analysis
79Llm

Model Collapse as Cultural Evolution

Dongxin Guo, Jikun Wu, Siu Ming Yiu

This paper argues that model collapse in LLMs is really a kind of cultural evolution problem, and shows that self-training can first improve then damage compositional language structure before eventually degrading it.

Read analysis