SDJul 2

UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation

arXiv:2607.016699.31 citations
Predicted impact top 35% in SD · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers working on text-to-music generation with limited data, this provides practical insights into batch sampling, though the gains are incremental.

This work investigates batch sampling strategies for text-to-audio music generation under low-data and small-scale model settings, finding that clustering training data by text embeddings outperforms audio embeddings, and that moderate cluster numbers optimize objective metrics while larger clusters improve structural coherence in listening tests.

This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describes our approach and findings for the ICME 2026 Grand Challenge on Academic Text-to-Music Generation. Training data are clustered using either text embeddings or audio embeddings, and samples with similar characteristics are grouped within the same mini-batch to mitigate gradient interference. The effects of modality and cluster granularity on clustering are analyzed. Results show that clustering based on text embeddings achieves better performance on objective evaluation metrics than clustering based on audio embeddings. In addition, different cluster granularity leads to different behaviors across evaluation criteria: a moderate number of clusters performs best on objective metrics, while a larger number of clusters tends to exhibit music with more coherent structure in listening tests.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes