Stanford
University
  • Stanford Home
  • Maps & Directions
  • Search Stanford
  • Emergency Info
  • Terms of Use
  • Privacy
  • Copyright
  • Trademarks
  • Non-Discrimination
  • Accessibility
© Stanford University.  Stanford, California 94305.
Skip to content
  • About

    • About
    • People
    • Get Involved with HAI
    • Support HAI
    • Subscribe to Email
  • Research

    • Research
    • Research Programs
    • Grants
    • Marlowe (opens in new tab)
    • Student Affinity Groups
    • Centers & Labs
    • Research Publications
    • Research Partners
  • Education

    • Education
    • Executive and Professional Education
    • Government and Policymakers
    • K-12
    • Stanford Students
  • Policy

    • Policy
    • Policy Publications
    • Policymaker Education
    • Student Opportunities
  • AI Index

    • AI Index
    • AI Index Report
    • Global Vibrancy Tool
    • People
  • News
  • Events
  • Industry
  • Centers & Labs
Navigate
  • About
  • Events
  • AI Glossary
  • Careers
  • Search
Participate
  • Get Involved
  • Support HAI
  • Contact Us

Stay Up To Date

Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.

Sign Up For Latest News

The Tests That Grade AI May Be Getting It Wrong | Stanford HAI
news

The Tests That Grade AI May Be Getting It Wrong

Date
September 25, 2026
Topics
Generative AI
Privacy, Safety, Security
Foundation Models

Benchmarks — the standardized tests that rank AI models on safety, bias, and reasoning — drive markets and shape regulation. New Stanford research finds they often don't measure what they claim to.

Before a new AI model reaches the public, its developers run it through a battery of tests known as “benchmarks,” which score it on everything from reasoning ability to how safe it is for people to use. Billions of investment dollars ride on these benchmark scores, and policymakers increasingly cite them to establish regulations and government procurement decisions that will guide the development of AI. 

In a pair of new studies to be presented in October at the Third Annual Conference on Language Modeling in San Francisco, Stanford researchers and their collaborators shifted the focus to the benchmarks themselves, asking a more basic question: Do these tests actually measure what they claim to? Based on their findings, they argue that the field needs to up its game and approach benchmarking with the seriousness it demands. 

Here, two of the researchers – Stanford Assistant Professor Sanmi Koyejo and graduate student Sang Truong – discuss their studies, which were partially supported by the Stanford Institute for Human-Centered AI (HAI), and what they mean for the future of AI benchmarking, and AI itself.

Let’s start with the basics. Who creates these benchmarks, and how do they work?

Koyejo: It used to be that benchmarks were mostly an academic exercise. By far the most famous benchmark is ImageNet, which Stanford HAI’s Fei-Fei Li built. Most public benchmarks are still built by academics, though companies and other stakeholders sometimes release them too. They’ve become a core part of the identity of AI. The benchmark creators are understandably interested in having their benchmarks become the standard. So, this question of whether benchmarks measure what they claim to is quite important. Borrowing tools from measurement science and psychometrics – the century-old science of measuring human abilities – our research shows that benchmarks do not always measure what they claim to, and that benchmarks claiming to measure the same thing often disagree with each other. In one of these studies, we ran that test across 56 widely used benchmarks and found the pattern repeatedly.

Truong: The AI developers adopt published benchmarks and use them to test their new model’s capability. How well it reasons, how safe it is, whether it’s biased toward or against certain people, and so on. Benchmarks are becoming a fundamental part of the AI industry to judge performance improvements within a given model or to compare performance across models. Their accuracy, reproducibility, and scalability are critical to AI’s future.

Can you give a real-world example of how a benchmark might fail to measure what it claims?

Truong: There’s a popular benchmark called BBQ that measures bias. It’s a series of multiple-choice questions. One family of BBQ questions gives you deliberately incomplete information. Something like: John and Mary are going to the gym; who is stronger? A) John; B) Mary; or C) We don’t know because we don’t have enough information. The correct answer is C, because you don’t know how old these people are or anything else about them. But if the model makes a leap in logic based on gender – that males tend to be stronger – and says “John,” then people running the test conclude the model has a gender bias.

Koyejo: But here’s the problem. A model that really is biased, if it’s also good at spotting a trick question, will answer “we don’t know” and score as unbiased. A model with no bias at all that simply misses the trick will score as biased. The question can’t discern those two apart. What it ends up measuring is closer to reading comprehension than to bias, and that’s a benchmark not measuring the thing its name promises. The benchmarks have to become more rigorous and scientific. That’s why we’ve turned to the long-standing fields of measurement science and psychometrics, which have gotten good at solving these challenges in the realm of human capabilities. 

How does measurement science solve these problems?

Truong: Measurement theory gives us a framework for how a score is produced and provides a statistical machinery that lets you test the hypotheses that two benchmarks measure what they claim to measure.

Koyejo: It’s the same theory behind standardized testing, like the SAT. There are many parallels between measuring AI models and assessing students’ learning. Psychometrics looks at this in two ways: convergent and discriminant validity. Convergent validity checks if benchmarks measuring the same property produce similar results. Discriminant validity checks if benchmarks measuring different properties actually show clear differences. Over the years, measurement science has gotten very good at making sure every test item precisely measures specific capabilities. We want the AI field to bring the same rigor to benchmarking.

In a second paper, you offer a compelling example of what can happen when benchmarks fall short, looking at how AI safety scores degrade in languages other than English. Can you explain this dynamic?

Truong: People worry a lot about AI safety, especially how easy it is to “jailbreak” a model – to get it to do something it’s specifically instructed not to do, like generating sexual or explicit content or providing bad advice in a mental health context. Safety benchmarks are incredibly important. Nearly all the major safety benchmarks are written in English, and when you translate them, two separate things break. The model’s guardrails might get weaker, so attacks that fail in English succeed in Spanish or Swahili. And the translated test becomes a shakier instrument, because the wording may drift or the question may get harder in translation. From a single score, you can’t tell which of those you’re looking at, and that’s exactly what measurement models let us pull apart. There are varying reasons for this: from too little of the non-English language in the training corpus to the prompt not being precisely translated or the translation making the question harder than it was intended to be. Measurement science lets us test which of these possibilities applies.

Koyejo: The most interesting aspect of applying measurement models here is that it lets us go from comparing a single number for “safety” to a much more nuanced understanding by breaking down safety performance into component factors, like language-agnostic safety robustness, the difficulty with which the model is being probed, general language processing difficulty, and how a particular prompt probes cross-lingual safety gaps. It’s really a much more detailed look at the many facets that comprise safety. Measurement science helps us get inside this complexity.

What’s at stake for the average person?

Truong: We argue that benchmarks are quite consequential to the future of AI. A model’s ranking on a performance leaderboard shapes its market value. It drives financial decisions by developers and investors, and it influences regulations by policymakers. Because benchmarks are key to data-driven, scientific decisions, we should get the numbers right – or at least know when they are wrong and how to interpret them. Measurement science can help with that.

Koyejo: I’d put it this way: Whoever holds the measuring stick steers the ship. Whatever AI is going to be, most of it will be shaped by the evaluation procedures, and if you get the evaluation wrong, the risks are quite great and rising. There’s also a public-good angle. As Sang mentioned, organizations and governments make procurement decisions against these numbers. Benchmarks are also starting to show up in regulation. I don’t think the question is whether they belong there. It’s whether anyone has checked that a particular score supports the particular decision being made from it. Usually, nobody has checked. That’s a fixable problem, not a reason to throw the numbers away. There are simple things we can do a lot better. Every field that measures well has put real effort into building and maintaining its instruments, but AI hasn’t done that yet. A benchmark score is a prediction about how a system will behave once it’s out in the world, and almost nobody records that prediction in a form that can be checked later against what actually happened. Until we do, we’re calibrating these instruments against each other and never against reality. That’s the missing piece, and it’s what we’re building.

Learn more:

  • Why Do Safety Guardrails Degrade Across Languages?

  • What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

Share
Link copied to clipboard!
Contributor(s)
Andrew Myers

Related News

Can AI Be Slowed Down? Stanford HAI Experts Weigh the Risks, Rules and Race Ahead
Shana Lynch
Sep 22, 2026
News

In a new series called Prompt Response, Stanford’s Surya Ganguli, Diyi Yang and Rob Reich examined emergent agent behavior, recursive self-improvement, independent evaluation and whether a kill switch can make advanced AI safer.

News

Can AI Be Slowed Down? Stanford HAI Experts Weigh the Risks, Rules and Race Ahead

Shana Lynch
Privacy, Safety, SecurityRegulation, Policy, GovernanceGenerative AISep 22

In a new series called Prompt Response, Stanford’s Surya Ganguli, Diyi Yang and Rob Reich examined emergent agent behavior, recursive self-improvement, independent evaluation and whether a kill switch can make advanced AI safer.

AI Chatbots And Privacy: How To Keep Personal Info Secure
Good Morning America
Sep 21, 2026
Media Mention

HAI Policy Fellow Jennifer King advises caution surrounding one's person security when using AI, citing her research on AI privacy and insights on users’ tendency to disclose personal information to chatbots due to their conversational nature, the lack of transparency in how companies use chatbot conversations for training, and the evolving legal landscape around this data.

Media Mention
Your browser does not support the video tag.

AI Chatbots And Privacy: How To Keep Personal Info Secure

Good Morning America
Privacy, Safety, SecuritySep 21

HAI Policy Fellow Jennifer King advises caution surrounding one's person security when using AI, citing her research on AI privacy and insights on users’ tendency to disclose personal information to chatbots due to their conversational nature, the lack of transparency in how companies use chatbot conversations for training, and the evolving legal landscape around this data.

Your Boss, Tech Companies And Police Can Read Your Chatbot Conversations
Washington Post
Aug 31, 2026
Media Mention

HAI Policy Fellow Jennifer King discusses privacy issues with AI chatbots, saying, “Unless you are having a chat with a service that has a temporary chat or, basically, an incognito version ... [and] you’re also having it within a browser that’s not tracking you, the answer is no. You can’t be sure that it’ll be totally private."

Media Mention
Your browser does not support the video tag.

Your Boss, Tech Companies And Police Can Read Your Chatbot Conversations

Washington Post
Privacy, Safety, SecurityAug 31

HAI Policy Fellow Jennifer King discusses privacy issues with AI chatbots, saying, “Unless you are having a chat with a service that has a temporary chat or, basically, an incognito version ... [and] you’re also having it within a browser that’s not tracking you, the answer is no. You can’t be sure that it’ll be totally private."