Danaé Metaxa

h-index1
7papers
1citation

7 Papers

9.1AIAug 3Code
MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models

Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian et al.

Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critical gap in LLM evaluation: no existing infrastructure routinely links interaction trajectories to user-defined outcomes. We introduce MonitrLLM, open-source infrastructure for community-centered LLM evaluations that links full conversation transcripts to user-reported task intent and outcome assessments, treating all three as primary evaluative signals rather than optional metadata. To demonstrate the value of this approach, we conducted a two-week feasibility pilot with 26 college students using ChatGPT, collecting 206 evaluation reports with full conversation transcripts. The findings from our pilot demonstrate the value of connecting conversation trajectories with user-reported outcomes. For instance, despite reporting high average satisfaction (4.19/5) with their LLM interactions, participants also experience a substantial 23.1% failure rate on their goal tasks. We also find that multi-turn conversations are reported as failing at 2.5 times the rate of single-turn exchanges, a pattern that reframes extended interaction as a signal of difficulty rather than engagement. We conclude by discussing the value of incorporating direct user feedback with observational data for robust LLM evaluations, and the possibilities for infrastructure that enables this goal.

6.5CYMay 10
Cost-of-Ethics Crisis: Beliefs, Decisions, and Justifications in the Job Searches of Computer Science Students in Canada and the United States

Mohamed Abdalla, Sahar Abdalla, Alicia Cappello et al.

Workplace norms in computer science have received growing attention due to a series of recent ethical scandals. One type of response has been a push to improve the ethics education provided to computer science students. Evidence for the effectiveness of ethics education remains mixed; some evidence suggests that norms are changing, while gaps between stated values and practice remain. Our focus here is on whether students, who have received some contemporary CS ethics education, are able to effectively apply ethical reasoning to their own decision-making in what is typically the first significant ethical decision of their careers: their job search. Our study examines the ethical decision making of 129 computer science students and recent graduates during their job searches. We find that most students prioritize factors like compensation, location, and workplace culture over ethical and social issues. Even when expressing ethical concerns, respondents often justify taking actions contradicting their moral views through commonly-shared explanations such as desire to make money or the perceived inability to avoid unethical workplaces. This work sheds light on the disconnect between ethics education and real-world CS graduate decision making. We offer insights for evolving curricula to better address practical ethical dilemmas, with implications for educators and industry.

6.6CYAug 6
The Algorithmic Flattening of Sound: Computational Evidence and Justice Implications of AI Music Homogenization

Zoe Slendebroek, Danaé Metaxa

This paper audits whether large-scale generative music systems exhibit measurable musical homogenization relative to human-produced music, and develops a justice-centered account of why this matters. We audit two commercially deployed systems (Suno and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, and Heavy Metal). For each system and genre, we generate 100 tracks and compare them against human corpora of equal size, using 72 music information retrieval (MIR) features and multiple diagnostics of dispersion, redundancy, and separability. We define homogenization as reduced acoustic variation in standard computational audio features including rhythm and timing, timbre/spectral shape, and dynamics, both within genres and across genre boundaries. We also generate tracks using only a genre name as the prompt, with no additional instructions, to reveal each system's default musical tendencies. The results show two structurally distinct homogenizing tendencies. Lyria reduces within-genre acoustic diversity, while Suno collapses the acoustic distinctions between genres without compressing within-genre spread. Neither system follows user prompts faithfully, indicating that the observed patterns reflect learned priors rather than prompt constraints. The two systems do not converge on a common acoustic profile and are more acoustically distant from each other than two random human subsamples would typically be. Nevertheless, a standard classifier distinguishes AI from human tracks near-perfectly on MIR features alone. We argue that these patterns matter not as an aesthetic curiosity but as a justice-relevant condition, shaping which musical styles become legible, valued, and economically rewarded as generated outputs increasingly circulate at scale.

10.9HCAug 6
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

Ro Encarnación, Tina Behzad, Emma Lurie et al.

Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT's chat UI and OpenAI's API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were less accurate than API responses on both benchmarks with search disabled. Enabling web search reduced accuracy by up to 8 percentage points, and even reversed the direction of modality performance trends for one benchmark. Repeated runs of the same prompt produced inconsistent responses in up to 21\% of prompts. The two modalities also grounded answers in different citations, and abstention behavior was also inconsistent across both modalities. These results illustrate that, even within a model family, reporting only simple accuracy metrics can obscure important forms of model behavioral variation relevant to AI safety assessments. We argue that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.

3.1CYAug 5
The Beginning of ChatGPT Ads

Emma Lurie, Ro Encarnación, Sorelle A. Friedler et al.

This paper presents the first empirical study of advertising content being rolled out in the user-facing online interfaces of large language models (LLMs). We systematically examine possible demographic differences in ad content shown to U.S. users of ChatGPT using a sock puppet audit methodology. We create and deploy 91 sock puppets in a 3x3 factorial design, using geolocation cues (account IP proxies and location-signaling prompts) to signal three racial/ethnic groups (Black, Hispanic, and White) and three income terciles (low, medium, and high). We conduct data collection starting in February 2026, collecting over 3,000 advertisements from 186 unique advertisers in response to 335 prompts on a range of realistic user queries. We find that accounts begin receiving ads 14 days after account creation, and that lower-income accounts, regardless of race, are more likely to receive ads. In this first phase of ChatGPT ads, the ads themselves skewed heavily towards consumer goods, directed users to a specific advertiser rather than a particular product, and were clearly separated from the LLM's response text, observations we anticipate will change as ads continue being integrated into LLM chat interfaces. We release a public, searchable archive of all collected advertisements. Finally, we discuss the implications of our findings, and conclude with methodological and theoretical recommendations for future empirical studies of LLM advertisements.

4.9CYJul 31
Triangulating Across U.S. Federal AI Transparency Regimes

Emma Lurie, Emma Fauser, Qing He et al.

Federal AI systems can deny benefits or flag individuals for deportation, but the public disclosures meant to make those systems visible are fragmented and unevenly detailed. This paper examines three existing U.S. federal transparency regimes---System of Records Notices (SORNs), Information Collection Requests (ICRs), and the AI Use Case Inventory---and asks how well they, individually and together, describe government AI use. We find that no single regime fully reveals how the government constructs or deploys AI: each discloses different aspects of a system, and the current disclosure infrastructure makes it very challenging for the public to track specific AI systems across regulatory regimes and over time. Persistent identifiers are absent, granularity varies widely, and the annual AI Use Case Inventory cycle means federal agencies can deploy systems months before appearing in any official record. Using hand-validated zero-shot classification and cross-document entity resolution, we contribute a triangulation method that links disclosures across all three regimes and present two case studies. Our case studies finds that linking records provides greater insight into government AI use, but even linked records would constitute insufficient oversight compared to what public reporting has revealed about the same systems. We trace each regime's disclosure weaknesses to its original administrative purpose, showing these gaps are structural, and offer recommendations focused on the AI Use Case Inventory as the mechanism best suited for public-facing transparency: (1) a broad and consistently applied AI system definition, (2) persistent system identifiers with cross-references to related disclosures, and (3) restored public visibility into risk management processes.

9.7HCJun 23
Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays

Megha N. Govindu, Stephanie T. Wang, Sorelle A. Friedler et al.

As large language models (LLMs) are increasingly used in media production from journalistm to filmmaking, what impact do they have on the stories being told? Prior work has shown LLMs to perpetuate social biases, including those related to gender. We complement existing literature on gender bias in LLM outputs by auditing the network structure of LLM-generated movie screenplays through automating the Bechdel test, a popular measure of women's representation in literary and film works. We also introduce the use of social network analysis measures to further analyze representational bias in LLM-generated scripts. We evaluate screenplays generated by three state-of-the-art LLMs (GPT-5, Gemini 3 Pro, and Claude Sonnet 4.5) against 768 corresponding human-written screenplays, finding that human-written scripts are more likely to pass the Bechdel test. However, other network analyses, like centrality, homophily, and triadic relationships demonstrate that in some cases LLM-scripts have less bias, although all script types demonstrate some representational bias under most measures. We conclude by discussing the continued need for further quantitative assessments of media representations and AI-generated content.