AdaTT: Text-Guided Instrument Timbre Transfer with Target-Adaptive Structural Control
For audio synthesis researchers, AdaTT improves timbre transfer quality by resolving conflicts between source instrument expressive details and target timbral properties.
AdaTT addresses timbral ambiguity in instrument timbre transfer by adaptively scaling pitch and loudness controls per frame based on text prompts, achieving superior timbral fidelity and naturalness while preserving score-level content.
This paper addresses timbral ambiguity in instrument timbre transfer under fine-grained structural conditions. We argue this issue stems from instrument-specific expressive details in these conditions, which conflict with the target timbral properties. For example, imposing a violin's pitch-dominant vibrato contours onto a flute, which naturally exhibits loudness-dominant vibrato, impairs timbral fidelity. We propose AdaTT, a target-adaptive system that ensures high timbral fidelity across diverse timbre transfer scenarios within the ControlNet scheme. It selectively scales the frame-wise influence of pitch and loudness controls via text prompts to match the target instrument's identity. We also present a semi-automatic data construction pipeline to teach the model which expressive details to transform or preserve. Results show AdaTT achieves superior timbral fidelity and naturalness while retaining score-level content. Audio samples are available at https://dabinkim0.github.io/adatt/.