10.9HCAug 6
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)Ro Encarnación, Tina Behzad, Emma Lurie et al.
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT's chat UI and OpenAI's API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were less accurate than API responses on both benchmarks with search disabled. Enabling web search reduced accuracy by up to 8 percentage points, and even reversed the direction of modality performance trends for one benchmark. Repeated runs of the same prompt produced inconsistent responses in up to 21\% of prompts. The two modalities also grounded answers in different citations, and abstention behavior was also inconsistent across both modalities. These results illustrate that, even within a model family, reporting only simple accuracy metrics can obscure important forms of model behavioral variation relevant to AI safety assessments. We argue that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.
3.1CYAug 5
The Beginning of ChatGPT AdsEmma Lurie, Ro Encarnación, Sorelle A. Friedler et al.
This paper presents the first empirical study of advertising content being rolled out in the user-facing online interfaces of large language models (LLMs). We systematically examine possible demographic differences in ad content shown to U.S. users of ChatGPT using a sock puppet audit methodology. We create and deploy 91 sock puppets in a 3x3 factorial design, using geolocation cues (account IP proxies and location-signaling prompts) to signal three racial/ethnic groups (Black, Hispanic, and White) and three income terciles (low, medium, and high). We conduct data collection starting in February 2026, collecting over 3,000 advertisements from 186 unique advertisers in response to 335 prompts on a range of realistic user queries. We find that accounts begin receiving ads 14 days after account creation, and that lower-income accounts, regardless of race, are more likely to receive ads. In this first phase of ChatGPT ads, the ads themselves skewed heavily towards consumer goods, directed users to a specific advertiser rather than a particular product, and were clearly separated from the LLM's response text, observations we anticipate will change as ads continue being integrated into LLM chat interfaces. We release a public, searchable archive of all collected advertisements. Finally, we discuss the implications of our findings, and conclude with methodological and theoretical recommendations for future empirical studies of LLM advertisements.
4.9CYJul 31
Triangulating Across U.S. Federal AI Transparency RegimesEmma Lurie, Emma Fauser, Qing He et al.
Federal AI systems can deny benefits or flag individuals for deportation, but the public disclosures meant to make those systems visible are fragmented and unevenly detailed. This paper examines three existing U.S. federal transparency regimes---System of Records Notices (SORNs), Information Collection Requests (ICRs), and the AI Use Case Inventory---and asks how well they, individually and together, describe government AI use. We find that no single regime fully reveals how the government constructs or deploys AI: each discloses different aspects of a system, and the current disclosure infrastructure makes it very challenging for the public to track specific AI systems across regulatory regimes and over time. Persistent identifiers are absent, granularity varies widely, and the annual AI Use Case Inventory cycle means federal agencies can deploy systems months before appearing in any official record. Using hand-validated zero-shot classification and cross-document entity resolution, we contribute a triangulation method that links disclosures across all three regimes and present two case studies. Our case studies finds that linking records provides greater insight into government AI use, but even linked records would constitute insufficient oversight compared to what public reporting has revealed about the same systems. We trace each regime's disclosure weaknesses to its original administrative purpose, showing these gaps are structural, and offer recommendations focused on the AI Use Case Inventory as the mechanism best suited for public-facing transparency: (1) a broad and consistently applied AI system definition, (2) persistent system identifiers with cross-references to related disclosures, and (3) restored public visibility into risk management processes.