CLAIJun 14

LLM-Assisted Stance Detection in Scientific Discourse: A Test Case in Bayesian Cognitive Science

arXiv:2606.1556613.7
Predicted impact top 75% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For social scientists conducting qualitative coding on interpretive, theoretically demanding constructs, this work provides a validated LLM-assisted framework that scales annotation while maintaining reliability, though it is a case study and not a general solution.

The authors developed an LLM-assisted stance detection method for a theoretically loaded construct (realism vs. instrumentalism in Bayesian cognitive science), achieving a held-out combined reliability of 0.76 and near-perfect article-level rank stability (r=0.96-0.97). They deployed it on 6,858 quotes from 210 articles, quantifying that low-level perception/motor articles scored 8.8 Realism points higher than high-level cognition articles.

Qualitative coding is central to social science, but expert annotation is difficult to scale. LLMs offer a possible extension, yet require careful validation when the target construct is interpretive, theoretically loaded, and only indirectly expressed. We study this problem in a difficult case: detecting whether authors treat Bayesian models as descriptions of mental and neural mechanisms (realism) or as useful mathematical tools (instrumentalism). Our method combines a theory-driven codebook, expert-coded reference annotations, a diagnostic-gated prompt-optimization search yielding a shared zero-shot prompt for three frontier LLMs (GPT-5.1, Claude Sonnet 4.6, Gemini 3 Pro Preview), and multi-rater reliability analysis. The final prompt achieved a held-out combined reliability score of 0.76 (harmonic mean of ICC = 0.79 and $α$ = 0.74), with all diagnostics satisfied. Deployed on 6,858 quotes from 210 articles, the three LLMs reached substantial quote-level agreement (ICC = 0.80; $α$ = 0.76; combined = 0.78) and near-perfect article-level rank stability ($r$ = 0.96-0.97 across rater pairs). The corpus was predominantly weakly realist, but article-level stances were rarely uniform: only 1.4% of articles used a single band, while 59.5% spanned four or more. Low-level perception/motor articles scored 8.8 Realism points higher than high-level cognition articles ($p < .001$, $d = 0.60$), quantifying a long-held qualitative intuition. We present this as an expert-led case study; the framework is intended to generalize to similar theoretically demanding tasks, not to all qualitative analysis.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes