One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation
AuthorsSanghyun Jo, Seo Jin Lee, Seohyung Hong, Yoorim Gang, Hyeongsub Kim, Hyungseok Seo, Kyungsu Kim
Resources
This paper turns dense cell segmentation from hundreds of clicks into one click per cell type by chaining prompts through SAM's built-in feature structure.
Key results
annotation cost reduction versus per-instance prompting on cell-type-annotated benchmarks
typical per-type prompts used by CoP on typed benchmarks
example performance reached with group prompting compared with per-instance upper bound
CoP with all components enabled on CoNIC
HSG only, before FPR on CoNIC
minimum precision maintained by Hierarchical Similarity Gating during iterations
What the paper found
This paper introduces Chain-of-Prompts, a training-free interactive framework built on the frozen image encoder of SAM3 that changes cell instance segmentation from one click per instance to one click per cell type. The key insight is that SAM3’s encoder already clusters same-type cells in feature space before prompting, so the method uses Hierarchical Similarity Gating to combine high-resolution and low-resolution cosine similarity maps, then applies non-parametric thresholding and connected-component labeling to extract reliable prompt points with precision above 96%. Farthest Prompt Recursion then selects the next prompt as the spatially farthest unreached reliable point, recursively expanding coverage without learned parameters. On three cell-type-annotated benchmarks—CoNIC, CoNSeP, and GlaS—CoP with only about 3 clicks per image retains over 90% of per-instance SAM3 performance while cutting annotation cost by 97% and even surpassing fully supervised models such as CellViT, CA-SAM2, and CellPose3. On four morphologically homogeneous benchmarks—MoNuSeg, TNBC, CryoNuSeg, and CPM-17—a single click per image preserves over 99% of per-instance prompting performance. The ablation study shows that removing recursive propagation drops AJI from 0.579 to 0.203, while the full method reaches 92.7% of the upper-bound AJI in an example with 3 clicks versus 245 per-instance clicks, demonstrating that group prompting is a practical way to scale foundation-model interaction to dense histopathology images.
Original abstract
Cell instance segmentation models trained on cell-specific datasets suffer severe performance drops on out-of-distribution cell types, while interactive foundation models overcome this through per-instance prompting at a cost that is prohibitively expensive for histopathology images containing hundreds to thousands of densely packed instances. We introduce Group Prompting, a new paradigm that shifts interactive segmentation from per-instance $O(N)$ to per-type $O(T)$, where a single click per cell type suffices to segment all instances of that type. Our key observation is that the frozen image encoder of the Segment Anything Model (SAM) already clusters same-type cells in its feature space before any prompt is given. Exploiting this property, we propose Chain-of-Prompts (CoP), a training-free framework that recursively expands a single user click by (1) identifying reliable same-type locations through non-parametric gating of multi-scale encoder features, and (2) selecting the most spatially distant reliable point as the next prompt to maximize coverage. On three cell-type-annotated benchmarks, CoP with one click per type retains over 90% of per-instance performance and surpasses fully-supervised methods without any additional training. On four morphologically homogeneous benchmarks, a single click retains over 99%. Project Page: https://shjo-april.github.io/Chain-of-Prompts/
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.