CL CVMar 14, 2024

Less is More: High-value Data Selection for Visual Instruction Tuning

Zikang Liu, Kun Zhou, Wayne Xin Zhao, Dawei Gao, Yaliang Li, Ji-Rong Wen

arXiv:2403.09559v44.85 citationsHas Code

Originality Incremental advance

AI Analysis

This addresses the problem of high training costs and inefficiency in building LVLMs for AI researchers and practitioners, offering an incremental improvement over existing data collection methods.

The paper tackles data redundancy in visual instruction tuning for large vision-language models by proposing TIVE, a high-value data selection method that uses gradient-based influence scores to select only about 15% of the data, achieving comparable or better performance on benchmarks.

Visual instruction tuning is the key to building large vision language models~(LVLMs), which can greatly improve the task generalization and solving capabilities by learning a mixture of instruction data from diverse visual tasks. Previous work mostly collects multiple existing visual instruction datasets via heuristic ways for training (even more than a million instructions), which may introduce data redundancy and enlarge the training cost. To investigate this issue, we conduct a series of empirical studies, which reveal a significant redundancy within the visual instruction datasets, and show that greatly reducing the amount of instructions from several tasks even do not affect the performance. Based on the findings, we propose a high-value data selection approach TIVE, to eliminate redundancy within the visual instruction data and reduce the training cost. In TIVE, we first estimate the instance influence score on its corresponding task, and the task difficulty score, based on the gradient-based influence functions. Then, we leverage the two kinds of scores to determine the task proportion within the selected visual instruction subset, and select high-value instances for each task, respectively. Experiments on various LVLMs show that our approach using only about 15% data can achieve comparable average performance to the full-data fine-tuned model across eight benchmarks, even surpassing it on four of the benchmarks. Our code and data will be publicly released.

View on arXiv PDF

Similar