The Benchmark Trap: Structures of Power and Injustice in AI Evaluations

| Source: arXiv AI

Tags: AI evaluation, benchmarks, AI ethics, leaderboards, structural injustice, research culture

A paper accepted to AIES 2026 argues AI benchmarks are not neutral evaluation tools but socio-technical systems that concentrate prestige and funding among powerful industry labs, applying Iris Marion Young's theory of structural injustice to benchmark culture.

Details

AI benchmarks determine who wins: who gets citations, trust, institutional influence, and funding. As compute costs rise, only well-resourced industrial labs can compete at the frontier, and benchmarks designed around SOTA performance systematically reward those labs regardless of broader scientific or social value. This paper, to appear at AIES 2026, frames these dynamics through Iris Marion Young's theory of oppression and structural injustice. The authors argue that benchmarking practices align with four of Young's five "faces of oppression" -- not because any actor behaves wrongly, but because individually defensible practices and network effects produce systematic harm when aggregated. The structural injustice framing is significant: the claim is not that researchers are malicious, but that the benchmarking ecosystem as designed perpetuates inequality and narrows viable research trajectories. The paper contends this prevents the field from advancing in epistemically robust and socially beneficial directions. The argument will resonate with those skeptical of leaderboard-driven AI development. It offers limited immediate practical recommendations but provides a useful theoretical vocabulary for critiques of benchmark-centric research culture.