CLJun 20

Are Multilingual Models Actually Improving? Isolating True Cross-Lingual Transfer

arXiv:2606.2195421.5
Predicted impact top 34% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For NLP researchers evaluating cross-lingual models, this provides a more accurate metric to measure transfer strength, correcting previous conflations.

The paper proposes a new metric, Hardness Adjusted Transfer (HAT) Score, to isolate true cross-lingual transfer from general accuracy improvements. Analysis across 20 models and 3 benchmarks reveals that small models transfer well, progress with model size is slower than expected, but clear progress has been made over time.

Cross-lingual transfer is a model's ability to generalize capabilities from well-represented source languages to under-represented target languages. Existing measures of a model's transfer strength conflate improvements in transfer with general improvements to accuracy in the source language. We advocate for an alternate metric that reliably captures transfer strength called Hardness Adjusted Transfer (HAT) Score, and use it to derive multiple insights on factors influencing transfer strength. Our analysis across twenty diverse language models and three popular mainstream multilingual benchmarks argues that 1) transfer in small models is not broken, 2) we are making slower than expected progress in cross-lingual transfer with model size, and 3) we have made clear progress over time.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes