CLGNECJun 29

Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments

arXiv:2606.3098716.9
Predicted impact top 42% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners needing scalable, interpretable evaluation of written expert judgments, EQMs provide a validated method that outperforms existing text-analysis and human ratings.

The authors introduce Explanation Quality Markers (EQMs), 60 reasoning patterns scored by LLMs, to measure judgment quality in natural-language explanations from forecasting tournaments. In over 55,000 forecast-rationale pairs, EQMs predict accuracy at both forecast and forecaster levels, outperforming pre-LLM methods, with over 90% of significant correlations matching hypotheses, and are the strongest predictor at the forecast level.

Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Forecasting tournaments offer a natural testing ground: probabilistic judgments are paired with natural-language rationales and scored against realized outcomes. We introduce Explanation Quality Markers (EQMs), a set of sixty theory-guided reasoning patterns scored by large language models (LLMs). In a pre-registered analysis of over 55,000 forecast-rationale pairs from a multiyear forecasting tournament, EQMs predict accuracy at both the forecast and forecaster levels, consistently outperforming pre-LLM text-analysis methods. More than 90% of statistically significant pattern-level EQM-accuracy correlations match our directional hypotheses. The signal is asymmetric: EQMs identify likely underperformers more reliably than they distinguish the very best forecasters. Benchmarked against traditional indicators of forecasting skill, EQMs are the strongest predictor at the forecast level and competitive at the forecaster level, though weaker than prior accuracy. Human ratings of rationale quality are less consistently correlated with accuracy and place disproportionate weight on rationale length. Results transfer to an independent forecasting study. EQMs provide a scalable, interpretable method for extracting judgment-relevant information from written explanations.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes