CYJun 26Code
Economic Evaluations of Language ModelsAlexander Wan, Stephane Hatgis-Kessell, Tomás Aguirre et al.
Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities relevant to tasks, work activities, and occupations in the US labor economy. We ground the evaluation suite in real user queries to language models where possible, and supplement these with synthetic data. Our evaluations improve coverage over OpenAI's GDPval benchmark, which is the existing state-of-the-art that covers 5% of US occupations, at 500x lower cost. Alongside benchmarks, we also introduce a simulation-based exposure measure to estimate how much time current language model capabilities could save across all tasks belonging to all US occupations, with detailed accounting for each estimate. Our estimates indicate that current models could save workers substantial time on at least half of their tasks in 47% of occupations. However, for 79% of tasks where we predict substantial time savings, observed Claude usage is low, suggesting that existing usage lags potential. Beyond inherent constraints of language model chatbots, our data identifies privacy and proprietary systems as the principal bottlenecks limiting further time savings from AI. Overall, we introduce adaptable infrastructure that grounds inferences about language models' labor-market impact in their current capabilities, which can be continually updated as capabilities improve.
CYJun 8
The Jagged Global Economy: Frontier AI Unevenly Exposes National EconomiesArul Murugan, Tomás Aguirre, Abhishek Nagaraj et al.
Frontier AI's labor-market effects matter to workers, firms, and policymakers, but current evidence generally comes from a handful of high-income economies. The capabilities of frontier AI are jagged across work tasks and national economies diverge in how they allocate human labor. We introduce a national AI exposure metric that combines occupation-level exposure scores and international employment data for 141 countries. We find that high income countries are substantially more exposed than low income countries and that Europe and Central Asia are 50 percent more exposed than Sub-Saharan Africa. We also find a gender gap: women are more exposed than men in 91 percent of countries, driven by their concentration in white-collar and sales occupations. The exceptions are countries where women's employment remains concentrated in agriculture and household enterprises. We validate our national AI exposure estimates by showing they predict national AI adoption statistics published by Anthropic, Microsoft, and OpenAI. Beyond direct exposure, we identify a new mechanism for indirect exposure due to cross-country income dependencies. Some nations such as Tajikistan depend heavily on foreign workers remitting money back to their home countries: Tajikistan's direct exposure to frontier AI is below-average but because 37 percent of Tajikistan GDP is Russian remittance and Russia is very exposed, Tajikistan's remittance-accounted exposure becomes above-average. Our research shows that national variation in exposure is large enough that policy responses calibrated to U.S. or European labor markets will not generalize.