GTCYJun 18

Impacts of Aggregation on Model Diversity and Consumer Utility

arXiv:2602.2329313.51 citationsh-index: 4
Predicted impact top 4% in GT · last 90 daysOriginality Incremental advance
AI Analysis

For AI marketplace designers and benchmark creators, this work identifies a flaw in current evaluation metrics and offers a theoretically grounded fix.

The paper shows that winrate, a standard LLM evaluation metric, incentivizes model homogenization and reduces consumer welfare, and proposes a weighted winrate mechanism that provably improves specialization incentives and consumer welfare.

Consider a marketplace of AI tools, each with slightly different strengths and weaknesses. By picking the right model for the task at hand, a user can do better than simply using the same model for everything. Routers operate under a similar principle, where sophisticated model selection can increase overall performance. However, aggregation is often noisy, reflecting imperfect user choices or routing decisions. This leads to two main questions: first, what does a "healthy marketplace" of models look like for maximizing consumer utility? Secondly, how can we incentivize producers to create such models? We show that winrate, a standard benchmark in LLM evaluation, can incentivize model creators to homogenize for both types of model changes, reducing consumer welfare. We propose a new mechanism, weighted winrate, which rewards models for answers that are higher quality, and show that it provably improves incentives for producers to specialize and increases consumer welfare. We conclude by exploring the impact of our theoretical results in empirical benchmark datasets and discussing implications for benchmark design.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes