CLLGMay 12, 2021

How Reliable are Model Diagnostics?

arXiv:2105.05641v131.6714 citations
Originality Synthesis-oriented
AI Analysis

This addresses the reliability of model diagnostics for researchers and practitioners, highlighting incremental concerns in evaluation methods.

The paper critically examines three recent diagnostic tests for pre-trained language models, finding that likelihood-based and representation-based diagnostics are not as reliable as assumed, and provides recommendations based on empirical findings.

In the pursuit of a deeper understanding of a model's behaviour, there is recent impetus for developing suites of probes aimed at diagnosing models beyond simple metrics like accuracy or BLEU. This paper takes a step back and asks an important and timely question: how reliable are these diagnostics in providing insight into models and training setups? We critically examine three recent diagnostic tests for pre-trained language models, and find that likelihood-based and representation-based model diagnostics are not yet as reliable as previously assumed. Based on our empirical findings, we also formulate recommendations for practitioners and researchers.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes