NTH

One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation

AuthorsSanghyun Jo, Seo Jin Lee, Seohyung Hong, Yoorim Gang, Hyeongsub Kim, Hyungseok Seo, Kyungsu Kim

July 5, 2026 2 min read
Watch on YouTube
The one-line take

This paper turns dense cell segmentation from hundreds of clicks into one click per cell type by chaining prompts through SAM's built-in feature structure.

Key results

97%
prompt reduction

annotation cost reduction versus per-instance prompting on cell-type-annotated benchmarks

3
clicks per example

typical per-type prompts used by CoP on typed benchmarks

92.7%
upper-bound AJI retention

example performance reached with group prompting compared with per-instance upper bound

0.579
AJI ablation full

CoP with all components enabled on CoNIC

0.203
AJI ablation without recursion

HSG only, before FPR on CoNIC

96%
precision

minimum precision maintained by Hierarchical Similarity Gating during iterations

What the paper found

This paper introduces Chain-of-Prompts, a training-free interactive framework built on the frozen image encoder of SAM3 that changes cell instance segmentation from one click per instance to one click per cell type. The key insight is that SAM3’s encoder already clusters same-type cells in feature space before prompting, so the method uses Hierarchical Similarity Gating to combine high-resolution and low-resolution cosine similarity maps, then applies non-parametric thresholding and connected-component labeling to extract reliable prompt points with precision above 96%. Farthest Prompt Recursion then selects the next prompt as the spatially farthest unreached reliable point, recursively expanding coverage without learned parameters. On three cell-type-annotated benchmarks—CoNIC, CoNSeP, and GlaS—CoP with only about 3 clicks per image retains over 90% of per-instance SAM3 performance while cutting annotation cost by 97% and even surpassing fully supervised models such as CellViT, CA-SAM2, and CellPose3. On four morphologically homogeneous benchmarks—MoNuSeg, TNBC, CryoNuSeg, and CPM-17—a single click per image preserves over 99% of per-instance prompting performance. The ablation study shows that removing recursive propagation drops AJI from 0.579 to 0.203, while the full method reaches 92.7% of the upper-bound AJI in an example with 3 clicks versus 245 per-instance clicks, demonstrating that group prompting is a practical way to scale foundation-model interaction to dense histopathology images.

Original abstract

Cell instance segmentation models trained on cell-specific datasets suffer severe performance drops on out-of-distribution cell types, while interactive foundation models overcome this through per-instance prompting at a cost that is prohibitively expensive for histopathology images containing hundreds to thousands of densely packed instances. We introduce Group Prompting, a new paradigm that shifts interactive segmentation from per-instance $O(N)$ to per-type $O(T)$, where a single click per cell type suffices to segment all instances of that type. Our key observation is that the frozen image encoder of the Segment Anything Model (SAM) already clusters same-type cells in its feature space before any prompt is given. Exploiting this property, we propose Chain-of-Prompts (CoP), a training-free framework that recursively expands a single user click by (1) identifying reliable same-type locations through non-parametric gating of multi-scale encoder features, and (2) selecting the most spatially distant reliable point as the next prompt to maximize coverage. On three cell-type-annotated benchmarks, CoP with one click per type retains over 90% of per-instance performance and surpasses fully-supervised methods without any additional training. On four morphologically homogeneous benchmarks, a single click retains over 99%. Project Page: https://shjo-april.github.io/Chain-of-Prompts/

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis