6.0CYMar 30
Judging the algorithm: Algorithmic accountability on the risk assessment tool for intimate partner violence in the Basque CountryAna Valdivia, Cari Hyde-Vaamonde, Julián García-Marcos
This paper discusses an algorithmic tool introduced in the Basque Country (Spain) to assess the risk of intimate partner violence. The algorithm was introduced to address the lack of human experts by automatically calculating the level of violence based on psychometric features such as controlling or violent behaviour. Given that critical literature on risk assessment tools for domestic violence mainly focuses on English-speaking countries, this paper offers an algorithmic accountability analysis in a non-English speaking region. It investigates the algorithmic risks, harms, and limitations associated with the Basque tool. We propose a transdisciplinary approach from a critical statistical and legal perspective. This approach unveils issues and limitations that could lead to unexpected consequences for individuals suffering from partner violence. Moreover, our analysis suggests that the algorithmic tool has a high error rate on severe cases, i.e., cases where the aggressor could murder his partner -- 5 out of 10 high-risk cases are misclassified as low risk -- and that there is a lack of appropriate legal guidelines for judges, the end users of this tool. The paper concludes that this risk assessment tool needs to be urgently evaluated by independent and transdisciplinary experts to better mitigate algorithmic harms in the context of intimate partner violence.
18.2HCApr 17, 2024
Characterizing and modeling harms from interactions with design patterns in AI interfacesLujain Ibrahim, Luc Rocher, Ana Valdivia · oxford
The proliferation of applications using artificial intelligence (AI) systems has led to a growing number of users interacting with these systems through sophisticated interfaces. Human-computer interaction research has long shown that interfaces shape both user behavior and user perception of technical capabilities and risks. Yet, practitioners and researchers evaluating the social and ethical risks of AI systems tend to overlook the impact of anthropomorphic, deceptive, and immersive interfaces on human-AI interactions. Here, we argue that design features of interfaces with adaptive AI systems can have cascading impacts, driven by feedback loops, which extend beyond those previously considered. We first conduct a scoping review of AI interface designs and their negative impact to extract salient themes of potentially harmful design patterns in AI interfaces. Then, we propose Design-Enhanced Control of AI systems (DECAI), a conceptual model to structure and facilitate impact assessments of AI interface designs. DECAI draws on principles from control systems theory -- a theory for the analysis and design of dynamic physical systems -- to dissect the role of the interface in human-AI systems. Through two case studies on recommendation systems and conversational language model systems, we show how DECAI can be used to evaluate AI interface designs.
4.9CLMay 20, 2025
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity TheoryFranziska Sofia Hafner, Ana Valdivia, Luc Rocher · oxford
Language models encode and subsequently perpetuate harmful gendered stereotypes. Research has succeeded in mitigating some of these harms, e.g. by dissociating non-gendered terms such as occupations from gendered terms such as 'woman' and 'man'. This approach, however, remains superficial given that associations are only one form of prejudice through which gendered harms arise. Critical scholarship on gender, such as gender performativity theory, emphasizes how harms often arise from the construction of gender itself, such as conflating gender with biological sex. In language models, these issues could lead to the erasure of transgender and gender diverse identities and cause harms in downstream applications, from misgendering users to misdiagnosing patients based on wrong assumptions about their anatomy. For FAccT research on gendered harms to go beyond superficial linguistic associations, we advocate for a broader definition of 'gender bias' in language models. We operationalize insights on the construction of gender through language from gender studies literature and then empirically test how 16 language models of different architectures, training datasets, and model sizes encode gender. We find that language models tend to encode gender as a binary category tied to biological sex, and that gendered terms that do not neatly fall into one of these binary categories are erased and pathologized. Finally, we show that larger models, which achieve better results on performance benchmarks, learn stronger associations between gender and sex, further reinforcing a narrow understanding of gender. Our findings lead us to call for a re-evaluation of how gendered harms in language models are defined and addressed.