CVLGJul 1

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation

arXiv:2607.0259220.1h-index: 1
Predicted impact top 8% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the limitation of static teacher routing in multimodal on-policy distillation, offering a more flexible token-level arbitration that improves reasoning quality for vision-language models.

H-OPD introduces a confidence-aware heterogeneous multi-teacher on-policy distillation framework for multimodal reasoning, achieving superior performance across 11 benchmarks by dynamically combining vision-language and text-only teachers at the token level.

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm by providing supervision on student-generated trajectories. However, existing OPD methods for multimodal reasoning usually rely on a static teacher routing, assigning each sample to a single teacher based on modality or task type. This ignores that visual grounding and abstract reasoning may dominate different decoding steps, making a single teacher insufficient for the full trajectory. To this end, H-OPD is proposed as a confidence-aware heterogeneous multi-teacher OPD framework for multimodal reasoning. By verifying the complementarity of heterogeneous teachers in the same reasoning process, H-OPD replaces task or sample level teacher routing with token-level teacher arbitration along the shared student trajectory. H-OPD employs vision-to-language description transfer to enable text-only teachers to access key visual semantics, and uses a confidence-aware arbitration mechanism to dynamically combine vision-language teacher and text-only teachers at each token. Extensive evaluations over 11 widely-used reasoning benchmarks showcase the superior performance of our method.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes