Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models
AuthorsSema Helali, Lina Abu Nada, Sausan Al Kawas, Alaa Abd-Alrazaq, Faleh Tamimi, Rafat Damseh
Resources
This paper maps the fast-growing landscape of AI in dentistry, showing how general-purpose and dental-specific foundation models can work together while highlighting key safety and data gaps.
Key results
DentVFM was pretrained on the DentVista corpus of about 1.6 million multimodal images
DentVLM was trained on 2.46 million visual question-answer pairs across 36 tasks
DentVLM reduced diagnostic time by 15–22% in the clinical evaluation
A detection-prompted tooth segmentation pipeline improved SAM Dice from 0.672 to 0.903
A contrastive CLIP variant achieved 0.929 accuracy for TMJOA diagnosis
What the paper found
This PRISMA-ScR scoping review synthesizes 97 studies from 2020–2026 on large AI models in dentistry and proposes a two-axis framework: architectural paradigm, separating language-generative models from discriminative vision foundation models, and degree of dental specialization, from general-purpose prompting to domain-specific pretraining. General-purpose systems, including OpenAI’s ChatGPT-4o and GPT-4, Anthropic’s Claude, Google’s Gemini, Microsoft Copilot, and DeepSeek, perform best on text-centric tasks such as licensing exams, clinical reasoning, and patient communication, with reported accuracies often reaching 80–95% on endodontic and guideline-based question sets, but they drop sharply on image-dependent tasks, frequently to 45–68% on radiographic questions. Vision models adapted from SAM, CLIP, and GroundingDINO achieve strong tooth segmentation and lesion detection only after dental fine-tuning, with one detection-prompted pipeline improving SAM Dice from 0.672 to 0.903 and a contrastive CLIP variant reaching 0.929 accuracy for TMJOA diagnosis. The strongest results come from dental-specific foundation models, especially DentVFM, pretrained on the DentVista corpus of about 1.6 million multimodal images, and DentVLM, trained on 2.46 million VQA pairs across 36 tasks; DentVLM outperformed junior dentists on 21 of 36 tasks and reduced diagnostic time by 15–22%. The review’s central finding is complementarity: structured pipelines that combine vision models for localization, retrieval-augmented generation, and generative models for explanation consistently outperform single-model systems. The main barriers to autonomous deployment remain hallucination, limited annotated dental data, and the absence of standardized benchmarks.
Original abstract
Background: Oral diseases affect nearly 3.5 billion people worldwide, yet the comparative clinical potential of large-scale AI models in dentistry remains poorly understood. Three distinct model categories have emerged: language-generative models, discriminative vision foundation models, and dental-specific foundation models, with no unified review examining their relationships and collective limitations. Methods: Following PRISMA-ScR guidelines, we systematically searched four databases (PubMed, Google Scholar, Scopus, arXiv), screened independently by two reviewers. After applying inclusion/exclusion criteria, 97 studies (2020-2026) were included. We propose a two-dimensional classification framework organizing models by architectural paradigm and dental specialization degree. Results: Language-generative models excel at text-based tasks (clinical reasoning, licensing exams, patient communication) but show inconsistent performance on image-dependent diagnostics. Adapted SAM and CLIP variants achieve strong tooth segmentation and lesion detection results. Dental-specific models (DentVFM, DentVLM, OralGPT) demonstrate strongest performance on complex multimodal tasks. Integrated pipelines consistently outperform single-model approaches. A data asymmetry is observed: dental-specific pretraining concentrates almost entirely in the vision domain, reflecting scarce large-scale dental text corpora. Conclusions: General-purpose and dental-specific models play complementary roles; the most effective systems combine both within structured pipelines. Safe autonomous deployment requires resolving three persistent barriers: hallucination in generative models, limited annotated dental datasets, and absent standardized clinical evaluation benchmarks.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.