CVGTJun 22

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

arXiv:2606.2291814.8
Predicted impact top 27% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers using VLMs as automated judges in video generation, this method improves evaluation accuracy by tailoring taxonomies to each model's capabilities.

JudgeFit discovers a per-VLM evaluation taxonomy for physical video consistency, outperforming a global-schema baseline on all 16 tested VLMs with a mean relative improvement of approximately 32%.

Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that provide reward signals, ranking decisions, and data-filtering criteria. Yet VLMs differ substantially in training data and architecture, encoding physical phenomena through distinct internal representations. A single global evaluation schema therefore gives every VLM the same axes of competence, regardless of what each can actually perceive. We propose JudgeFit, an iterative refinement procedure that discovers a per-VLM evaluation taxonomy. An initial taxonomy is constructed by prompting the target VLM to enumerate physics errors on a small set of videos and clustering the resulting descriptions. The taxonomy is then refined through a diagnostic step: we calibrate the VLM's per-dimension scores to human physical-commonsense ratings, diagnose which dimensions it scores unreliably or redundantly, and prompt an LLM to repair them, iterating until convergence. We further instantiate this procedure as a benchmark and apply it to 16 VLMs spanning eight model families. The refined taxonomy outperforms the global-schema baseline on held-out videos for every VLM tested, with a mean relative improvement of approximately 32%. Beyond aggregate accuracy, the per-VLM profiles expose model-specific blind spots that overall rankings cannot anticipate, with reliability patterns differing markedly across model families.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes