Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

arXiv:2607.1340818.4h-index: 18
Predicted impact top 6% in AS · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners in text-to-audio generation, this work provides a method to improve instruction following for multi-event and temporal order tasks, which is a known bottleneck in current models.

The paper addresses the problem of text-to-audio models failing to follow instructions involving multiple sound events and temporal order. They propose using audio-aware large language models as fine-grained judges to provide feedback for direct preference optimization, achieving improvements in event completeness, temporal ordering, and joint instruction-following accuracy while maintaining audio quality.

Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limited supervision on instruction-level correctness. We propose an instruction-level framework that uses audio-aware large language models (ALLMs) as fine-grained judges to verify target event presence and temporal relations in generated audio. After validating ALLM judgments on benchmarks and through human verification, we use their feedback to construct preference pairs for direct preference optimization. We further introduce S3Bench, a narrative benchmark for evaluating multi-event temporal instruction following. Experiments show that our method improves event completeness, temporal ordering, and joint instruction-following accuracy across existing benchmarks and S3Bench, while maintaining audio quality.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes