NTH

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

AuthorsShravan Venkatraman, Omkar Thawakar, Ritesh Thawkar, Abdelrahman Shaker, Rao Muhammad Anwer

August 15, 2026 2 min read
Watch on YouTube
The one-line take

CVPD helps multimodal models notice visual details they can perceive but previously failed to use, without relying on external teachers or annotations.

Key results

15000
Unlabeled image pool

Images processed for counterfactual blind-spot discovery.

2590
Curated training tuples

Self-discovered image, question, region, and answer tuples.

17.2%
Curation yield

Fraction of processed images producing a retained blind spot.

3.60
OCRBench improvement

Point gain for Qwen3-VL-8B-Instruct.

3.38
MMStar Fine-Grained Perception improvement

Point gain for Qwen3-VL-8B-Instruct.

3.08
MMStar Logical Reasoning improvement

Point gain for Qwen3-VL-8B-Instruct.

What the paper found

CVPD, or Contrastive Counterfactual Visual Process Distillation, introduces a self-contained way to improve multimodal large language models without labels, reward models, segmentation tools, or stronger annotators. Using Qwen3-VL-8B-Instruct, it generates fine-grained questions and probes image regions through three counterfactual views: the full image, an enlarged crop, and a Gaussian-blurred “ghost” image. A three-gate criterion selects visual blind spots where the crop changes and sharpens the answer distribution, while removing the region leaves the model’s default behavior nearly unchanged. The crop becomes a positive teacher, the ghost becomes a negative teacher, and a contrastive per-token objective transfers latent perceptual ability into the full-image policy while KL anchoring preserves general capabilities. From 15,000 unlabeled images, CVPD creates 2590 curated training tuples, a 17.2% yield. On Qwen3-VL-8B-Instruct, it outperforms six self-evolving baselines across twelve benchmarks, including approaches using OpenAI’s GPT-4o during data construction, without regression. Gains are largest on localized perception: OCRBench improves by 3.60 points, MMStar Fine-Grained Perception by 3.38 points, and MMStar Logical Reasoning by 3.08 points. The results suggest that visual improvement can come from teaching a model to exploit perceptual information it already encodes, rather than adding external supervision.

Original abstract

Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce \textbf{CVPD} (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning. We propose a three-gate Counterfactual Criterion that identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation. On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression. It achieves gains of $+3.60$ on OCRBench, $+3.38$ on MMStar Fine-Grained Perception, and $+3.08$ on MMStar Logical Reasoning, while maintaining or improving performance on broader multimodal benchmarks.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis