Can LLMs Hire Fairly? Racial Bias in Resume Screening
For policymakers and employers using LLMs in hiring, this reveals that newer models may have reversed the direction of racial bias, but the underlying fairness remains unaddressed.
This paper audits 14 LLMs for racial bias in resume screening, finding that the sole 2023 model reproduces a pro-White callback gap (+2.12 pp), while all 2024+ models show null or pro-Black gaps (up to -3.01 pp), indicating a reversal in algorithmic hiring bias across generations.
We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (2022). The sole 2023-vintage model reproduces the pro-White callback gap documented in field experiments on labor market discrimination ($+2.12$ pp, significant at the 1\% level). Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal (up to $-3.01$ pp). The same pattern holds on the gender axis. Based on 24,024 paired postings per model across 14 models, our results document a reversal in the direction of algorithmic hiring bias across model generations.