CVJun 21

Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition

arXiv:2606.224168.6
Predicted impact top 59% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For video action recognition with long-tailed data, Gen2Balance provides a practical generative balancing method that significantly boosts rare-class accuracy.

Gen2Balance uses a text-to-video generative model to balance long-tailed video datasets, improving accuracy by 5.1% on UCF-LT and 7.0% on K100-LT, and by 31.9% on rare actions in RareAct.

We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video generative model, conditioned on diverse text prompts grounded in action profiles and training exemplars. Our approach, called Gen2Balance, converts an imbalanced training set into a balanced combination of real and generated video clips. To effectively learn from such data, we employ a two-stage training strategy that mitigates domain shift and yields significant improvements. We evaluate on long-tailed versions of standard benchmarks: UCF-101 (UCF-LT) and a 100-class subset of Kinetics (K100-LT) selected to prioritise temporally challenging actions. Gen2Balance improves accuracy over the strongest baselines for long-tailed learning by 5.1% and 7.0% on the respective datasets. On rare actions from the RareAct dataset (e.g., cut keyboard), Gen2Balance improves accuracy by 31.9%, demonstrating effectiveness for scarce actions. By varying the amount of synthetic data added, we show that partial balancing already achieves 79% of the performance gains at 27% of the compute cost on K100-LT, highlighting the practical scalability of Gen2Balance.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes