NTH
Research collection

Foundation Models research

Explore broadly trained models and their adaptation to new tasks and domains. Follow research on scaling, transfer, and evaluation.

47 papers · Latest edition October 7, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Foundation Models papers

Newest editions first.

01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis
04Foundation Model

nnFoundation: 3D Foundation Models for Radiology

Constantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, Bálint Kovács, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica Kächele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Krämer, Maximilian Zenk, Tim Rädsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David Kügler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina ...

nnFoundation trains complementary 3D convolutional and transformer models on 2.1 million medical scans and shows that matching architecture to the task can improve transfer across diverse radiology applications.

Read analysis
05Foundation Model

$t_0$: A Time-Series Foundation Model for Forecasting with Context

Lucas Meyer, Claudio Sole, Huikan Xiang, Nicolas Li, Lucas Franceschino, Arnau Quera-Bofarull, Maarten P. Scholl, Joachim Fainberg, Geoffrey Négiar

t0 is an open-weight time-series foundation model that uses historical and future context to deliver strong zero-shot probabilistic forecasts across diverse datasets.

Read analysis
06Foundation Model

TabPFN-3.5: Technical Report

Benjamin Jäger, Nick Erickson, Léo Grinsztajn, Felix Birkel, Klemens Flöge, Oscar Key, Kürşat Kaya, Jonas Kübler, Adèle Frankel, Tobias Schröder, Anurag Garg, Jan Hendrik Metzen, David Salinas, Simon Bing, Kristina Collins, Tuana Çelik, Vahid Balazadeh, Lydia Sidhoum, Tomás Pereda, Brendan Roof, Andrej Tschalzev, Siyuan Guo, Philipp Singer, Lennart Purucker, Jake Robertson, Marie Salmon, Philipp Jund, Jerry Chen, Diana Kriuchkova, Arthur Cahu, Eliott Kalfon, Adrian Hayler, Georg Grab, Vitor Monteiro, Lilly Wehrhahn, Dominik Safaric, Clara Cornu, Alan Arazi, Rylee Grace, Simone Alessi, Mihir Manium, Bernhard Schölkopf, Yann LeCun, Madelon Hulsebos, Sauraj Gambhir, Noah Hollmann, Frank Hutter

TabPFN-3.5 advances tabular foundation models with stronger multimodal and real-world data handling, faster inference, and scalable reasoning modes.

Read analysis
07Foundation Model

AI for Games in the Foundation Model Era

Meng Luo, Yanlin Li, Hao Li, Hongzhan Lin, Pengfei Zhou, Tianjie Ju, Ran Zhang, Yeying Jin, Mong-Li Lee, Wynne Hsu

The paper maps how foundation models are transforming game playing, design, development, adaptation, and testing while emphasizing that game-specific validation remains essential.

Read analysis
09Foundation Model

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Yuxuan Ren

ZGCM-1 shows how a fully open 7B model can combine efficient training, long-context reasoning, and tool-using agents to compete with much larger systems.

Read analysis
10Foundation Model

Learning Human Health and Diseases from 24-hour Wrist Movement

Yong Wang, Dylan McGagh, Katya Broomberg, Zizheng Zhang, Jonathan Carter, Junayed Naushad, Laura Brocklebank, Yang Sun, George Nicholson, Dianjianyi Sun, Canqing Yu, Jun Lv, Maxim Barnard, Hubert Lam, Andrew Steptoe, David W. Eyre, Liming Li, Zhengming Chen, Naomi Wray, Spiros Denaxas, Gary S. Collins, Huaidong Du, Aiden Doherty, Hang Yuan

Sensori turns a day of ordinary wrist movement into a reusable health representation that can improve prediction of many diseases without requiring clinical visits or retraining.

Read analysis
12Foundation Model

Xiaomi-TabLDM: A Tabular Foundation Model Technical Report

Xiaomi-TabLDM Team, :, Penghui Wang, Wei Liu, Hong Wang, Chengyue Huang, Yuxi Sun, Zirui Wang, Hongming Huang, Quan Wang, Chunxiao Liu, Erli Meng, Bin Wang

Xiaomi-TabLDM is a synthetic-data-trained tabular foundation model that aims to deliver strong, efficient classification and regression without task-specific fine-tuning.

Read analysis
13Foundation Model

EXAONE Tabular 1.0 : Technical Report

Moonjung Eo, Min-Kook Suh, Hye-Seung Cho, Jiwon Kim, Seoyoon Kim, Sangjun Nam, Soonyoung Lee

EXAONE Tabular shows that a compact, causally synthetic-pretrained foundation model can deliver strong tabular predictions without fine-tuning, at far lower cost than much larger alternatives.

Read analysis
14Foundation Model

OneModel: A Unified Foundation for Platform-Scale Multi-Scenario Ranking

Yinqi Zhang, Peiyu Hu, Yuntian Tang, Siying Gu, Jiahao Liang, Longxin Kou, Haiqing Hu, Shuman Zhuang, Yubin Xu, Chenggen Sun, Bin Ye, Donghui Xu, Zhaoyu Liu, Jiang Rong, Yuting Jia, Zhaokai Luo, Leilei Ma, Yiying Xie, Yao Hu

OneModel unifies user behavior across feeds, ads, and commerce into one production-scale ranking system that improves engagement and business metrics.

Read analysis
16Foundation Model

Forecast Collapse in Time-Series Foundation Models

Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu, Stephen Wang, Qingsong Wen, Huan Liu

This paper shows that time-series models can make deceptively flat forecasts and introduces a method that improves stock ranking without sacrificing calibration.

Read analysis
17Foundation Model

Human-Centric Intelligence in the Era of Foundation Models: A Survey

Yang Chen, Tianqi Wang, Xiaorui Jiang, Yilei Man, Yihua Shao, Mengyuan Liu, Zhi Chen, Xiaofeng Cao, Qibin Zhao, Chi Harold Liu, Albert Y. Zomaya, Nicu Sebe, Jingren Zhou, Dacheng Tao, Song Guo, Jingcai Guo

This survey maps how foundation models can evolve from recognizing people to understanding their actions, interactions, environments, and embodied agency.

Read analysis
19Foundation Model

K-EXAONE 2.0 Technical Report

Eunbi Choi, Kibong Choi, Sehyun Chun, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Ahra Jo, Hyunjik Jo, Yeonsik Jo, Minhyeok Jung, Doyoung Kim, Heegyu Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil Kim, Byungoh Ko, Changhun Lee, Dohaeng Lee, Haeju Lee, Jinsik Lee, Kyungmin Lee, Minwoo Lee, Wonkee Lee, Sangha Park, Sungjune Park, Kwangrok Ryoo, Kijung Seo, Minju Seo, Yongwoo Song, Sejong Yang, Heuiyeen Yeen, Stanley Jungkyu Choi, Yemuk Choi, Yongchan Chun, Jiwon Ham, Dasol Hong, Sujeong Im, Kijeong Jeon, Gerrard Jeongwon Jo, Hyeongjun Jo, Yujin Jo, Jiyeon Jung, Naeun Kang, Daeseong Kim, Euisoon Kim, Hayeon Kim, Hyosang Kim, Myoungshin Kim, Unsol Kim, Youchul Kim, Chaeeun Lee, ChaeYoon Lee, Edward Hwayoung Lee, Honglak Lee, Hwansoo Lee, Minkyung Lee, Sangeun Lee, Solji Lim, Woohyung Lim, Chanwoo Moon, Jueun Mun, Jimin Park, Seojeong Park, Yongmin Park, Hyerin Seo, Donghyeon Shin, Donghyun Son, Eunyong Son, Kaehyun Um, Sihoon Yang, Chang En Yea, Sihyuk Yi, K...

K-EXAONE 2.0 is a 750B-parameter open-weight multilingual MoE model designed for long-context reasoning, coding agents, and culturally grounded safety.

Read analysis
22Foundation Model

Bridging Compute- and Data-Optimal Pretraining

Tian Qin, Kimia Hamidieh, David Alvarez-Melis

This paper helps answer how to spend compute wisely when high-quality training data is scarce by modeling the diminishing value of repeated and paraphrased tokens.

Read analysis
23Foundation Model

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen, Yanru Chen, Yifei Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Dazhi Cheng, Yean Cheng, Jialei Cui, Jingbing Cui, Anqi Dai, Jiaqi Deng, Hao Ding, Rui Ding, Shaofeng Ding, Mengfan Dong, Mengnan Dong, Yuhao Dong, Yuxin Dong, Angang Du, Chenzhuang Du, Dikang Du, Jusen Du, Yulun Du, Yu Fan, Jing Feng, Qiulin Feng, Yichen Feng, Kelin Fu, Qiang Fu, Fuxuan Gao, Hongcheng Gao, Jingyue Gao, Tong Gao, Weijia Gao, Shangyi Geng, Jie Gong, Linhu Gong, Shengao Gong, Xiaochen Gong, Qizheng Gu, Yicheng Gu, Shuhao Guan, Haiqing Guo, Shiqi Guo, Xiang Guo, Zhengyan Guo, Beixi Hao, Wenxin Hao, Xiaoru Hao, Dailan He, Haotian He, Lehan He, Qi He, Weiran He, Xinran He, Xinyi ...

Kimi K3 is an open 2.8-trillion-parameter multimodal model designed to deliver frontier-level reasoning, coding, and long-horizon agent capabilities at unprecedented scale.

Read analysis
24Foundation Model

TiRex-2: Generalizing TiRex to Multivariate Data and Streaming

Patrick Podest, Marco Pichler, Elias Bürger, Levente Zólyomi, Bernhard Voggenberger, Wilhelm Berghammer, Daniel Klotz, Sebastian Böck, Günter Klambauer, Sepp Hochreiter

TiRex-2 is a recurrent time-series foundation model designed to forecast many interacting variables continuously without repeatedly recomputing the entire history.

Read analysis
25Foundation Model

Loop the Loopies!

Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai

Loopie argues that repeatedly looping a sparse MoE Transformer can beat simply scaling model size, reaching remarkable mathematical and scientific reasoning performance.

Read analysis
26Foundation Model

A Sovereign, Open-Source Foundation Model for German and English

The Soofi-Team, :, Benedikt Droste, David Fitzek, Ruben Härle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom Röhr, Sebastian von Rohrscheidt, Jörg Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim Köhler, Alexander Löser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max Lübbering

Soofi S is a new open bilingual foundation model for German and English that combines MoE and Mamba-Transformer ideas to deliver strong quality with efficient long-context inference.

Read analysis
27Foundation Model

Scalable Visual Pretraining for Language Intelligence

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen

This paper argues that training on the visual form of documents, not just extracted text, can improve foundation models by preserving information like layouts, equations, and figures.

Read analysis
28Foundation Model

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

Zhengpeng Feng, Sadiq Jaffer, Ira Shokar, Jovana Knezevic, Mark Elvers, Clement Atzberger, Robin Young, Aneesh Naik, Niall Robinson, Andrew Blake, David Coomes, Anil Madhavapeddy, Srinivasan Keshav

TESSERA v2 shows how to scale Earth-observation foundation models more effectively, revealing that downstream results—not pretraining loss—should guide model selection and that bigger encoders plus distillation can produce compact, highly competitive embeddings.

Read analysis
29Foundation Model

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Wentao Zhang, Liliana Hotsko, Woojeong Kim, Pengyu Nie, Stuart Shieber, Yuntian Deng

Program-as-Weights lets a model compile a natural-language task description into a small reusable neural program that can run locally and cheaply instead of calling a giant LLM every time.

Read analysis
31Foundation Model

The State-Prediction Separation Hypothesis

Giovanni Monea, Nathan Godey, Kianté Brantley, Yoav Artzi

This paper argues that transformers should split remembering from predicting, and that doing so can make language models train more efficiently and work better downstream.

Read analysis
34Foundation Model

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Lianghua Huang, Zhifan Wu, Wei Wang, Yupeng Shi, Mengyang Feng, Junjie He, Chenwei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Zoubin Bi

Wan-Streamer is a single multimodal foundation model built for real-time two-way audio-visual conversation with sub-second latency, replacing many separate pipeline modules.

Read analysis
36Foundation Model

HRM-Text: Efficient Pretraining Beyond Scaling

Guan Wang, Changling Liu, Chenyu Wang, Cai Zhou, Yuhao Sun, Yifei Wu, Shuai Zhen, Luca Scimeca, Yasin Abbasi Yadkori

HRM-Text claims you can pretrain a capable language model far more cheaply by swapping Transformers for a brain-inspired recurrent design and training on instruction pairs instead of internet-scale raw text.

Read analysis
39Foundation Model

NITP: Next Implicit Token Prediction for LLM Pre-training

Xiangdong Zhang, Debing Zhang, Shaofeng Zhang, Xiaohan Qin, Yu Cheng, Junchi Yan

NITP tweaks LLM pretraining by making models predict not just the next token, but its hidden semantic representation too, improving downstream performance with little extra cost.

Read analysis
40Foundation Model

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

Mind Lab, :, Song Cao, Vic Cao, Kaijie Chen, Bunny Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Hongquan Gu, Aaron Guan, Mutian Hong, Hailee Hou, Peixuan Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Autumn Jin, Fancy Kong, Kyrie Lei, Alexy Li, Dawn Li, Ray Li, Theo Li, Wenhao Li, Jiayi Lin, Domini Liu, Heshan Liu, Kairus Liu, Logan Liu, Maeve Luo, Runism Lv, Pony Ma, Verity Niu, Anson Qiu, Vincent Wang, Maxwell Yao, Regis Ye, Wenlin Ye, Yanying Ye, Josh Ying, Danney Zeng, Salmon Zhan, Anya Zhang, Ruijia Zhang, Shiyang Zhang, Sueky Zhang, Ya Zhang, Wei Zhao, Ada Zhou, Sizer Zhou, Xinyue Zhu, Murphy Zhuang

This paper argues that small adapters on top of foundation models could become persistent personal models at massive scale, enabling individualized behavior without retraining huge models.

Read analysis
41Foundation Model

Unified Neural Scaling Laws

Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish

This paper proposes a single scaling-law formula that better predicts how neural networks behave as model size, data, compute, and other factors all change at once.

Read analysis
42Foundation Model

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

Jing Huang, Daniel Wurgaft, Rachit Bansal, Laura Ruis, Naomi Saphra, David Alvarez-Melis, Andrew Kyle Lampinen, Christopher Potts, Ekdeep Singh Lubana

The paper explains why bigger models can learn rare, complex tasks that smaller ones miss: they have enough capacity to avoid overwriting those features while learning the common ones.

Read analysis
43Foundation Model

LLMSurgeon: Diagnosing Data Mixture of Large Language Models

Yaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Jiacheng Liu, Xinyue Bi, Zhaoyi Li, Zhiqiang Shen

LLMSurgeon tries to infer what data an LLM was trained on just by reading its outputs, offering a new way to audit foundation models without access to their training sets.

Read analysis
45Foundation Model

A Mechanistic Study of Tabular Foundation Models

Marin Biloš, James T. Wilson, Anderson Schneider, Yuriy Nevmyvaka

This paper peels back the internals of tabular foundation models to show how different architectures make predictions, why they become permutation-invariant, and how they fail under targeted attacks.

Read analysis
46Foundation Model

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation

Zhengrui Guo, Zhengyu Zhang, Jiabo Ma, Yihui Wang, Fengtao Zhou, Yingxue Xu, Ling Liang, Chenglong Zhao, Qi Xie, Jinbang Li, Shujing Guo, Fangyi Han, Zhijian Cen, Ziyi Liu, Cheng Jin, Junlin Hou, Zhixuan Chen, Yu Cai, Lijuan Qu, Shifu Chen, Yueping Liu, Zhe Wang, Xiuming Zhang, Muyan Cai, Li Liang, Hao Chen

This work introduces a clinically validated lung pathology foundation model that not only performs well across many diagnostic tasks, but also improves pathologist accuracy, speed, and consistency in real-world trials.

Read analysis
47Foundation Model

Towards a General Intelligence and Interface for Wearable Health Data

Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff

This paper turns massive wearable sensor data into a foundation model that can predict health outcomes, estimate daily metrics, and power a personalized health assistant.

Read analysis