NTH

Human-Centric Intelligence in the Era of Foundation Models: A Survey

AuthorsYang Chen, Tianqi Wang, Xiaorui Jiang, Yilei Man, Yihua Shao, Mengyuan Liu, Zhi Chen, Xiaofeng Cao, Qibin Zhao, Chi Harold Liu, Albert Y. Zomaya, Nicu Sebe, Jingren Zhou, Dacheng Tao, Song Guo, Jingcai Guo

August 22, 2026 2 min read
Watch on YouTube
The one-line take

This survey maps how foundation models can evolve from recognizing people to understanding their actions, interactions, environments, and embodied agency.

Key results

6
Human context levels

Taxonomy levels from visual appearance through embodied agency.

300M
Sapiens training data

Human images used for foundation-scale dense human perception.

1B
Sapiens2 training data

Human images used to expand dense perception and model scale.

8.4K
GR00T N1 training data

Hours of data used for NVIDIA’s generalist humanoid policy.

2.2B
GR00T N1 model size

Parameters in the generalist humanoid control model.

What the paper found

This survey reframes human-centric intelligence as a connected foundation-model discipline rather than a collection of isolated vision, motion, interaction, and robotics tasks. Its central taxonomy spans 6 levels: visual appearance, spatial geometry, kinematic dynamics, interaction modeling, world simulation, and embodied agency, progressing from observable subjects to dynamic actors and situated agents. Methodologically, it organizes systems into single-step mapping, autoregressive sequential factorization, iterative generation with diffusion or flow matching, and hybrid architectures that combine reasoning, generation, and control. It also compares training strategies including full fine-tuning, LoRA-based parameter-efficient tuning, instruction tuning, reward optimization, distillation, and inference methods such as retrieval augmentation and guided sampling. Representative scaling trends include Sapiens, trained on 300M human images, and Sapiens2, expanded to 1B images for dense pose, parsing, depth, and surface-normal prediction. In motion and embodiment, MotionGPT and LLaMo treat human movement as a language-compatible modality, while world models such as EgoSim connect action-conditioned prediction to planning. Human experience is increasingly used for robot learning: NVIDIA’s GR00T N1 uses 8.4K hours of data and a 2.2B-parameter policy, while GPT-4 exemplifies the role of general-purpose vision-language reasoning in humanoid control. The survey’s main conclusion is that scale alone is insufficient: future systems must align multimodal human representations with metric geometry, physical constraints, causal world states, privacy-aware data, and closed-loop execution, evaluated through integrated benchmarks rather than task-specific scores alone.

Original abstract

Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, and general-purpose modeling. Yet it has not fully integrated with foundation models to achieve the comparable progress seen in them. More importantly, recent advances across this broad landscape remain fragmented across tasks, modalities, and research communities, leaving their intrinsic conceptual and methodological connections unclear. To bridge these divides and rethink human-centric intelligence in the foundation-model era, we introduce a full-spectrum human context taxonomy that integrates six interconnected levels by viewing humans as observable subjects through visual appearance and spatial geometry, as dynamic actors through kinematic dynamics and interaction modeling, and as situated agents through world simulation and embodied agency. We next present the methodological foundations of the field, covering human-centric data families, computational architecture paradigms, and representative training and inference optimization strategies. We then systematically review representative methods across these levels and organize the associated datasets, benchmarks, and evaluation metrics. We further discuss open challenges and promising research directions toward human-centric intelligence that is scalable, trustworthy, physically grounded, and deployable, aiming to provide a coherent framework and practical reference for advancing the field. Finally, we provide a systematically organized and continuously updated collection of human-centric AI literature and resources on our project page.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis