NTH

Towards a General Intelligence and Interface for Wearable Health Data

AuthorsGirish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff

May 22, 2026 3 min read
Watch on YouTube
The one-line take

This paper turns massive wearable sensor data into a foundation model that can predict health outcomes, estimate daily metrics, and power a personalized health assistant.

Key results

100K to 100M parameters
Model scaling range

The paper studies joint scaling of model capacity across SensorFM variants from XXSmall to Base.

31%
Reconstruction loss reduction

At the largest 5M-subject scale, the biggest SensorFM variant achieves a 31% lower reconstruction validation loss than the smallest model.

74.8%
Random imputation improvement

SensorFM outperforms the best baselines by 74.8% on the random-imputation generative task, showing strong missing-data recovery.

ΔAUC = 0.09
Downstream task gains

Across discriminative classification tasks, scaling SensorFM yields a mean AUC improvement of 0.09 over smaller models.

Δr = 0.21
Downstream task gains

Across regression tasks, scaling SensorFM yields a mean Pearson-correlation improvement of 0.21 over smaller models.

35
Clinical tasks

SensorFM is evaluated on 35 person-level clinical and behavioral prediction tasks spanning cardiovascular, metabolic, mental health, sleep, lifestyle, and demographics.

What the paper found

This paper introduces SensorFM, a wearable-health foundation model pretrained with a masked autoencoder objective on more than 1 trillion minutes of unlabeled sensor data from 5 million participants spanning five modalities: PPG-derived heart rate and HRV, accelerometry, electrodermal activity, skin temperature, altitude, and sleep signals. Scaling both data and model size from 100K to 100M parameters produced near-linear gains, including a 31% lower reconstruction loss at the largest scale, a 74.8% reduction in random-imputation MSE, and downstream gains of ΔAUC = 0.09 on classification and Δr = 0.21 on regression. On 35 person-level clinical tasks across cardiovascular, metabolic, mental health, sleep, lifestyle, and demographic endpoints, a frozen SensorFM encoder with simple linear probes outperformed engineered-feature baselines on 34 of 35 tasks and achieved the best score on 31 of 35, with larger models also reducing dependence on demographic priors. The model’s generative capability enabled high-fidelity recovery of missing wearables: after ablating 60 contiguous minutes, it preserved 99.7% step-count accuracy, 99.9% deep-sleep accuracy, and 99.2% light-exercise accuracy in daily summaries. The authors then used a Gemini-based “classroom” of LLM agents to autonomously search downstream heads, running over 30,516 experiments and improving performance on 29 of 35 tasks versus linear probing. Finally, integrating SensorFM into a Personal Health Agent improved clinician-rated context, personalization, relevance, justifiability, and harm scores across 1,860 blinded ratings, with no significant difference versus using ground-truth labels, showing that model-derived predictions can ground safer, more useful health dialogues.

Original abstract

While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging. Specifically, converting low-level sensor data into representations capable of characterizing higher-level states is difficult due to high phenotypic diversity and variation in individual baseline health, physiology, and lifestyle factors. Moreover, collecting wearable data paired with health outcome annotations is laborious and expensive, and retrospective annotation remains practically unfeasible, contributing to a scarcity of data with high-quality labels. To overcome these limitations, we propose a foundation model for wearable health that is pretrained on more than one trillion minutes of unlabeled sensor signals drawn from a large cohort of five million participants. We demonstrate that the joint scaling of model capacity and pretraining data volume leads to systematic improvements in performance, as evaluated on a diverse set of 35 health prediction tasks, spanning cardiovascular, metabolic, sleep, and mental health, as well as lifestyle choices and demographic factors. We find that this population scale representation unlocks label-efficient few-shot learning and generative capabilities for robust daily metric estimation. To further leverage this learned representation, we deploy a classroom of LLM agents to autonomously search the space of downstream predictive heads built on the model embeddings, showing broad performance improvements that increase with LLM model capacity. Finally, we show how integrating these downstream predictors into a Personal Health Agent can support model responses that are more relevant, contextually aware, and safe, and we validate this via 1,860 ratings from a cohort of clinicians.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis