SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection
For researchers and practitioners using LLMs for stance detection, SICI provides a diagnostic tool to understand and predict model failures, revealing that high-complexity examples remain a bottleneck not resolved by common interventions.
The authors introduce SICI, a seven-dimensional diagnostic measure of semantic-pragmatic complexity for stance detection, which predicts LLM accuracy better than surface proxies and reveals regime shifts in error patterns across complexity levels, with high-complexity examples concentrating on None predictions. The measure shows substantial cross-scorer reliability (α=0.771) and the phase-transition structure persists across multiple LLMs.
Prompt-based LLMs are increasingly used for stance detection, but harder examples are not always repaired by clearer instructions, reasoning prompts, retrieval, or debate. We introduce SICI (Stance Inference Complexity Index), a seven-dimensional diagnostic measure of the semantic-pragmatic burden imposed by a target--text pair. Across SemEval-2016 and VAST, SICI predicts LLM accuracy better than surface proxies and shows substantial cross-scorer reliability ($α=0.771$). More importantly, LLM errors change regime as SICI increases: low-complexity examples invite over-attribution, especially Against predictions; intermediate examples form an unstable boundary; and high-complexity examples rapidly concentrate on None. This phase-transition-like structure persists across GPT-3.5, GPT-4o-mini, DeepSeek-V3, and GPT-4o, although stronger models move the boundaries. A 15-method intervention study further shows that prompting, retrieval, and debate often shift models along the attribution--abstention axis rather than removing the high-complexity bottleneck.