Stanford
University
  • Stanford Home
  • Maps & Directions
  • Search Stanford
  • Emergency Info
  • Terms of Use
  • Privacy
  • Copyright
  • Trademarks
  • Non-Discrimination
  • Accessibility
© Stanford University.  Stanford, California 94305.
NeurIPS 2025 Paper: Fantastic Bugs and Where to Find Them in AI Benchmarks | Stanford HAI
Skip to content
  • About

    • About
    • People
    • Get Involved with HAI
    • Support HAI
    • Subscribe to Email
  • Research

    • Research
    • Research Programs
    • Grants
    • Marlowe (opens in new tab)
    • Student Affinity Groups
    • Centers & Labs
    • Research Publications
    • Research Partners
  • Education

    • Education
    • Executive and Professional Education
    • Government and Policymakers
    • K-12
    • Stanford Students
  • Policy

    • Policy
    • Policy Publications
    • Policymaker Education
    • Student Opportunities
  • AI Index

    • AI Index
    • AI Index Report
    • Global Vibrancy Tool
    • People
  • News
  • Events
  • Industry
  • Centers & Labs

Stay Up To Date

Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.

Sign Up For Latest News

Navigate
  • About
  • Events
  • AI Glossary
  • Careers
  • Search
Participate
  • Get Involved
  • Support HAI
  • Contact Us
news

NeurIPS 2025 Paper: Fantastic Bugs and Where to Find Them in AI Benchmarks

Date
December 04, 2025
Topics
Communications, Media

By Sang Truong

AI benchmarks shape the trajectory of AI development. Our evaluations are only as good as the questions we ask. Right now, many benchmark questions are not doing so well.

In our NeurIPS 2025 paper, Fantastic Bugs and Where to Find Them in AI Benchmarks, we introduce an efficient way to detect flawed benchmark items at scale.

We identify three major categories of problematic questions: 

  • Ambiguous questions

  • Incorrect answer keys

  • Grading issues (for example: the correct answer is “4” but the grader marks “4.00” as incorrect)

Manually auditing benchmarks is very costly. MMLU alone spans 57 domains and contains 14,000 questions. Using the core assumption of unidimensionality in evaluation, we propose three measurement-theoretic statistics that automatically flag problematic items for expert review.

Across nine widely used benchmarks, our framework helps human experts identify flawed questions with up to 84% precision. The paper contains many interesting examples.

If you are at NeurIPS in San Diego, please visit our poster 1403 on Friday, Dec 5, 2025, from 11 AM to 2 PM

This work is joint with Yuheng Tu, Michael Hardy, Anka Reuel, Zeyu Tang, Jonathan Perera, Chibuike Uwakwe, Ben Domingue, Nick Haber, and Sanmi Koyejo

Reposted from LinkedIn with the author's permission.


This article is a part of the Stanford Data Science legacy publication. Read more about the HAI and Stanford Data Science merger.

Share
Link copied to clipboard!

Related News

Reading Today’s Headlines Through AI: A Real-Time Audit of Six Commercial Chatbots
Mirac Suzgun and James Zou
Jun 03, 2026
News

In a new study, scholars measured how accurately popular AI chatbots answered questions about the emerging news and found substantial regional disparity, dependence on distinct information ecosystems, and acute fragility under imperfect prompts.

News

Reading Today’s Headlines Through AI: A Real-Time Audit of Six Commercial Chatbots

Mirac Suzgun and James Zou
Communications, MediaGenerative AIJun 03

In a new study, scholars measured how accurately popular AI chatbots answered questions about the emerging news and found substantial regional disparity, dependence on distinct information ecosystems, and acute fragility under imperfect prompts.

CORES Faculty Director Featured in "The Transmitter": A brief history of precision self-scanning
Jan 21, 2026
Media Mention
Media Mention

CORES Faculty Director Featured in "The Transmitter": A brief history of precision self-scanning

Communications, MediaJan 21
Marlowe Computing Spotlight: Andreas Tolias Lab
Jan 20, 2026
Media Mention
Media Mention

Marlowe Computing Spotlight: Andreas Tolias Lab

Communications, MediaJan 20