Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots
AuthorsShravan Venkatraman, Omkar Thawakar, Ritesh Thawkar, Abdelrahman Shaker, Rao Muhammad Anwer
Resources
CVPD helps multimodal models notice visual details they can perceive but previously failed to use, without relying on external teachers or annotations.
Key results
Images processed for counterfactual blind-spot discovery.
Self-discovered image, question, region, and answer tuples.
Fraction of processed images producing a retained blind spot.
Point gain for Qwen3-VL-8B-Instruct.
Point gain for Qwen3-VL-8B-Instruct.
Point gain for Qwen3-VL-8B-Instruct.
What the paper found
CVPD, or Contrastive Counterfactual Visual Process Distillation, introduces a self-contained way to improve multimodal large language models without labels, reward models, segmentation tools, or stronger annotators. Using Qwen3-VL-8B-Instruct, it generates fine-grained questions and probes image regions through three counterfactual views: the full image, an enlarged crop, and a Gaussian-blurred “ghost” image. A three-gate criterion selects visual blind spots where the crop changes and sharpens the answer distribution, while removing the region leaves the model’s default behavior nearly unchanged. The crop becomes a positive teacher, the ghost becomes a negative teacher, and a contrastive per-token objective transfers latent perceptual ability into the full-image policy while KL anchoring preserves general capabilities. From 15,000 unlabeled images, CVPD creates 2590 curated training tuples, a 17.2% yield. On Qwen3-VL-8B-Instruct, it outperforms six self-evolving baselines across twelve benchmarks, including approaches using OpenAI’s GPT-4o during data construction, without regression. Gains are largest on localized perception: OCRBench improves by 3.60 points, MMStar Fine-Grained Perception by 3.38 points, and MMStar Logical Reasoning by 3.08 points. The results suggest that visual improvement can come from teaching a model to exploit perceptual information it already encodes, rather than adding external supervision.
Original abstract
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce \textbf{CVPD} (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning. We propose a three-gate Counterfactual Criterion that identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation. On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression. It achieves gains of $+3.60$ on OCRBench, $+3.38$ on MMStar Fine-Grained Perception, and $+3.08$ on MMStar Logical Reasoning, while maintaining or improving performance on broader multimodal benchmarks.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.