CVAIJul 7

LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding

arXiv:2607.0576910.0
Predicted impact top 40% in CV · last 90 daysOriginality Highly original
AI Analysis

For researchers and practitioners in optical music recognition and document understanding, Legato 2 provides a more accurate and comprehensive method for extracting both notation and textual content from sheet music.

Legato 2 introduces the first OMR model that processes sheet music system-by-system and generates symbolic transcriptions with embedded text, achieving new state-of-the-art performance across multiple datasets.

We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music. Legato 2 features the first large-scale neural model for optical music recognition (OMR) to operate sequentially on a system-by-system basis, following the horizontal lines of notation as they are read on the page, rather than treating the page as an undifferentiated image, enabling better scaling to arbitrarily long inputs. It is also the first OMR model capable of generating symbolic transcriptions that include embedded textual content, such as titles and annotations. The pipeline combines system-level segmentation with an autoregressive vision-LM to capture both local notation details and score structure. Across multiple datasets, Legato 2 consistently outperforms prior state of the art. We also show that symbolic transcriptions complement visual inputs for frontier language models, improving their interpretation of dense musical documents. Legato 2 establishes new state-of-the-art performance in both OMR and downstream sheet music understanding.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes