CLJun 16

Montreal Forced Aligner and the state of speech-to-text alignment in 2026

arXiv:2606.1846615.3Has Code
Predicted impact top 65% in CL · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers and practitioners needing accurate forced alignment, this paper presents an updated, widely-used tool with strong performance, though improvements are incremental.

The paper documents the development of MFA 3.0 over the past decade and evaluates its performance, achieving state-of-the-art or near state-of-the-art results with mean boundary errors below 15 ms across English, Japanese, and Korean benchmarks.

The Montreal Forced Aligner (MFA) was released in 2016 and has since become the most widely used tool for forced alignment in research and industry. In the decade since, MFA has undergone substantial development, including expanded coverage across more languages and dialects using larger open-source datasets, harmonized IPA dictionaries, model adaptation, cross-language phone remapping, and support utilities. This paper documents MFA 3.0's developments since version 1.0 and evaluates MFA's performance across English, Japanese, and Korean, benchmarked against classic and neural forced aligners. MFA 3.0 achieves state-of-the-art or near state-of-the-art performance across all four benchmark datasets with mean boundary errors below 15 ms. Adaptation and cross-language remapping are effective for languages outside MFA's training distribution, and pronunciation probability modeling and phonological rules provide gains in specific conditions.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes