NTH

Lethe: How Hard Is It to Forget? A Benchmark for Federated Unlearning in Medical Imaging

AuthorsShengchao Chen, Ting Shu

August 5, 2026 2 min read
Watch on YouTube
The one-line take

Lethe tests whether federated medical-AI models can truly forget hospitals, patients, or classes—and finds that difficult forgetting requests, rather than the choice of method, reveal the real weaknesses.

Key results

12
Unlearning methods

Methods evaluated in the Lethe benchmark.

8
Task families

Clinical task families spanning classification, segmentation, synthesis, registration, denoising, localization, and VQA.

16
Datasets

Datasets included across the benchmark.

0.572
FedRecovery BloodMNIST MIA

Calibrated membership-inference AUC, where ideal performance is 0.5.

107.2x
FedDNI BloodMNIST speedup

Unlearning-step speedup relative to full retraining.

What the paper found

Lethe, from Shengchao Chen at the University of Technology Sydney and Ting Shu at Shenzhen University, is a benchmark for federated unlearning in medical imaging, evaluating 12 methods across 8 task families and 16 datasets, including MedMNIST, Kvasir-SEG, IXI MRI, and medical VQA, with client-, class-, and sample-level forgetting. Using retraining on retained data as the gold standard, the benchmark measures utility, privacy, backdoor erasure, efficiency, model closeness, and durability. Its central finding is that forgetting difficulty, rather than the unlearning algorithm, determines whether methods can be distinguished: ordinary client removal often produces nearly identical task performance because medical models generalize across hospitals, while hard class-level or sole-class removal exposes meaningful differences. On class-level forgetting, Prune and FedQUIT preserve retain accuracy close to retraining, whereas gradient ascent can collapse utility; even accuracy-matched methods may retain forget-class representations. Consequently, residual membership—not task accuracy—is often the key privacy signal: on BloodMNIST, FedRecovery reaches a calibrated membership-inference AUC of 0.572, compared with the ideal 0.5. Most methods are far cheaper than retraining; FedDNI achieves a 107.2x speedup on BloodMNIST, although FedEraser’s dependence on stored update history makes it slower than retraining in some settings. Lethe recommends difficulty-aware evaluation, membership-focused auditing, and deletion-ready architectures such as site-scoped adapters rather than relying solely on post-hoc weight editing.

Original abstract

Federated learning enables medical-imaging models to be trained across hospitals, and privacy law, most explicitly the GDPR ``right to be forgotten'', turns removing a hospital's, a class's, or a patient's influence from such a model into a federated unlearning problem. This need is most acute in medicine, where patients withdraw consent and hospitals leave collaborations. Yet nearly all unlearning evidence comes from natural images, whose heterogeneity and task structure differ sharply from clinical data, so it is unclear whether existing methods transfer, and no shared protocol covers clinical data. We present Lethe, a benchmark for federated unlearning in medical imaging. It evaluates twelve methods across eight task families, from classification and segmentation to denoising, cross-modality synthesis, and vision-language question answering, at three forgetting granularities and against a retrained gold standard on utility, privacy, and cost. The central result is that what separates methods is the difficulty of the forgetting request, not the method itself. The easy removals that dominate the literature leave the methods that preserve utility indistinguishable, while only hard ones separate them. More striking, on the many medical tasks that generalize across sites, forgetting a client barely changes task performance, leaving residual membership as the signal that must be erased.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis