SDLGJul 9

MuScriptor: An Open Model for Multi-Instrument Music Transcription

arXiv:2607.0816811.3h-index: 13
Predicted impact top 19% in SD · last 90 daysOriginality Incremental advance
AI Analysis

It addresses the lack of robust, open-source multi-instrument transcription models for real-world music, though it is an incremental improvement over existing methods.

MuScriptor is an open-weight model for multi-instrument music transcription that works on real-world recordings across diverse genres, achieved by combining synthetic pre-training, fine-tuning on real audio, and reinforcement learning post-training.

Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes