What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
AuthorsMeera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang
Resources
This study asks whether popular AI benchmarks measure what they claim, finding that many safety and capability tests may not reliably distinguish the concepts they are intended to evaluate.
Key results
Capability and safety benchmarks initially evaluated.
Instruction-tuned models spanning major open and closed model families.
Benchmarks remaining after saturation and format-compliance exclusions.
Predictive effect of shared score format on benchmark-ranking similarity.
Predictive effect of shared assigned concept on benchmark-ranking similarity.
Evidence that BBQ-accuracy aligns more with reasoning than bias.
What the paper found
This paper tests whether AI benchmarks measure the concepts they claim to measure by adapting convergent and discriminant validity from psychometrics. The researchers evaluated 56 capability and safety benchmarks across 53 instruction-tuned models, including OpenAI’s GPT-4, GPT-5, o1, Anthropic’s Claude, Meta’s Llama, DeepSeek, Microsoft’s Phi, and xAI’s Grok. Using Spearman correlations between model rankings and one-parameter logistic item-response theory, they found that safety benchmarks labeled refusal, safety detection, and bias often converge weakly, while capability benchmarks labeled reasoning and knowledge correlate so strongly across categories that they frequently fail to discriminate between them. After removing saturated or format-sensitive tests, 48 benchmarks remained for the main analysis. Benchmark design often mattered more than the intended construct: a partial Mantel analysis found that shared score format predicted similarity with β = 0.275, compared with β = 0.138 for shared concept, and LLM-judge scoring was especially influential. The individual-benchmark analysis also exposed likely construct mismatches: BBQ-accuracy, commonly used in commercial releases from OpenAI, Anthropic, and Google DeepMind, aligned more strongly with reasoning than bias, with a relabeling statistic of 0.15, while DecodingTrust-Fair aligned more with knowledge. In contrast, refusal and over-refusal benchmarks were strongly inversely related, indicating genuine conceptual separation. The released item- and benchmark-level dataset enables future evaluation audits rather than treating leaderboard scores as direct measures of capability or safety.
Original abstract
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt convergent and discriminant validity from the social sciences into an approach for interrogating AI benchmarks, applying it to 56 capability and safety benchmarks across 53 models. We label benchmarks with substantively similar purported concepts to a shared assigned concept, and ask whether model rankings on benchmarks with the same assigned concept correlate more strongly than rankings on benchmarks with different assigned concepts. We ask analogous questions at the item level using item response theory (IRT) models. We find that correlations between model rankings on benchmarks with the same assigned safety concepts are often weak, suggesting these concepts may be conceptualized inconsistently across benchmarks. For assigned capability concepts (e.g., reasoning, knowledge), model rankings are often as strongly correlated among benchmarks with the same assigned concept as between benchmarks with different assigned concepts, suggesting these capability concepts may not discriminate well from one another. In some cases, benchmarks that share design elements (e.g., score format) correlate more strongly than benchmarks with the same assigned concept. Finally, some individual benchmarks correlate more strongly with benchmarks assigned a different concept than with benchmarks sharing their own assigned concept, suggesting they may measure a different concept than they purport to. For example, BBQ-accuracy correlates more strongly with benchmarks labeled reasoning than with benchmarks that share its assigned concept, bias. To support future empirical work on benchmark validity, we release our extensive dataset of model outputs and scores at the item- and benchmark-level.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.