Stanford
University
  • Stanford Home
  • Maps & Directions
  • Search Stanford
  • Emergency Info
  • Terms of Use
  • Privacy
  • Copyright
  • Trademarks
  • Non-Discrimination
  • Accessibility
© Stanford University.  Stanford, California 94305.
Sanmi Koyejo | Beyond Benchmarks: Building a Science of AI Measurement | Stanford HAI
Skip to content
  • About

    • About
    • People
    • Get Involved with HAI
    • Support HAI
    • Subscribe to Email
  • Research

    • Research
    • Research Programs
    • Grants
    • Marlowe (opens in new tab)
    • Student Affinity Groups
    • Centers & Labs
    • Research Publications
    • Research Partners
  • Education

    • Education
    • Executive and Professional Education
    • Government and Policymakers
    • K-12
    • Stanford Students
  • Policy

    • Policy
    • Policy Publications
    • Policymaker Education
    • Student Opportunities
  • AI Index

    • AI Index
    • AI Index Report
    • Global Vibrancy Tool
    • People
  • News
  • Events
  • Industry
  • Centers & Labs
Navigate
  • About
  • Events
  • AI Glossary
  • Careers
  • Search
Participate
  • Get Involved
  • Support HAI
  • Contact Us

Stay Up To Date

Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.

Sign Up For Latest News

Your browser does not support the video tag.
eventSeminar

Sanmi Koyejo | Beyond Benchmarks: Building a Science of AI Measurement

Status
Past
Date
Wednesday, March 19, 2025 12:00 PM - 1:15 PM PST/PDT
Location
Gates Computer Science Building Room 119
Topics
Sciences (Social, Health, Biological, Physical)
Attend Virtually

The widepread deployment of AI systems in critical domains demands more rigorous approaches to evaluating their capabilities and safety.

Share
Link copied to clipboard!
Event Contact
Annie Benisch
abenisch@stanford.edu

Related Events

Confronting Our AI Future: Hope, Fear, and the Choices Ahead
ConferenceOct 28, 20269:00 AM - 6:45 PM
October
28
2026

The rapid acceleration of AI comes with a profound wave of anxiety. Across every sector of society, people are facing unsettling questions about their worth and their place in a shifting world.

Conference

Confronting Our AI Future: Hope, Fear, and the Choices Ahead

Oct 28, 20269:00 AM - 6:45 PM

The rapid acceleration of AI comes with a profound wave of anxiety. Across every sector of society, people are facing unsettling questions about their worth and their place in a shifting world.

Tim de Silva | AI Financial Advice: Supply, Demand, and Life Cycle Implications
SeminarOct 07, 202612:00 PM - 1:15 PM
October
07
2026
Seminar

Tim de Silva | AI Financial Advice: Supply, Demand, and Life Cycle Implications

Oct 07, 202612:00 PM - 1:15 PM
NVIDIA & Marlowe | GPU Computing Foundations - Amanda Butler
Oct 07, 2026
October
07
2026

This session covers the foundational knowledge of GPUs, including their architecture, functionality, and applications in computing. It provides an introduction to GPU computing through the lens of the Marlowe SuperPod and prepares learners for advanced topics such as GPU-accelerated data science and machine learning.

Location: CoDa W401

Event

NVIDIA & Marlowe | GPU Computing Foundations - Amanda Butler

Oct 07, 2026

This session covers the foundational knowledge of GPUs, including their architecture, functionality, and applications in computing. It provides an introduction to GPU computing through the lens of the Marlowe SuperPod and prepares learners for advanced topics such as GPU-accelerated data science and machine learning.

Location: CoDa W401

While current evaluation practices rely on static benchmarks, these methods face fundamental efficiency, reliability, and real-world relevance challenges. This talk presents a path toward a measurement framework that bridges established psychometric principles with modern AI evaluation needs. We demonstrate how techniques from Item Response Theory, amortized computation, and predictability analysis can substantially improve the rigor and efficiency of AI evaluation. Through case studies in safety assessment and capability measurement, we show how this approach can enable more reliable, scalable, and meaningful evaluation of AI systems. This work points toward a broader vision: evolving AI evaluation from a collection of benchmarks into a rigorous measurement science that can effectively guide research, deployment, and policy decisions.

Speaker
Sanmi Koyejo
Assistant Professor of Computer Science, Stanford University; Faculty Affiliate, Stanford HAI

Watch the Event Recording