A Two-Stage Statistical Framework for Evaluating Associative Interference in Large Language Models
For researchers evaluating bias in LLMs, this provides a principled framework to isolate associative interference from refusal behavior, addressing a key methodological confound.
This paper adapts the Implicit Association Test to a forced-choice framework and introduces a two-stage model to separate response compliance from task-consistent classification in LLMs. Across three models, Claude Sonnet-4 showed strong interference in Gender-Career (DeltaP=0.086), while GPT-5 exhibited minimal interference, demonstrating that associative asymmetries are model-specific and can be mitigated.
Large language models (LLMs) are increasingly evaluated for bias using adaptations of human psychological paradigms, yet methodological limitations-particularly the conflation of refusal behavior with task performance-have hindered clear interpretation. Here, we adapt the Implicit Association Test (IAT) to a controlled, forced-choice framework and introduce a two-stage modeling approach that separates response compliance from task-consistent classification. Across three contemporary LLMs (Claude Sonnet-4, Gemini 2.5 Pro, and GPT-5), we evaluate associative interference, defined as reduced task-consistency in incongruent relative to congruent conditions. While compliance with the structured response format was uniformly high, interference effects varied substantially across models and domains. Claude Sonnet-4 exhibited strong interference in the Gender--Career domain (DeltaP = 0.086, 95% CrI [0.026, 0.173]) and smaller but credible effects in Gender--Science. Gemini 2.5 Pro showed attenuated interference, and GPT-5 exhibited minimal or no detectable interference across domains. These findings demonstrate that IAT-style associative asymmetries are not a universal property of LLMs, but instead depend on model-specific characteristics. By isolating interference from compliance and modeling item-level variability, this study provides a principled framework for evaluating structured response patterns in LLMs. The results highlight the importance of model-specific assessment and suggest that associative interference can be substantially mitigated in modern systems.