Sanmi Koyejo | Beyond Benchmarks: Building a Science of AI Measurement
The widepread deployment of AI systems in critical domains demands more rigorous approaches to evaluating their capabilities and safety.
Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.
Sign Up For Latest News
The widepread deployment of AI systems in critical domains demands more rigorous approaches to evaluating their capabilities and safety.
The rapid acceleration of AI comes with a profound wave of anxiety. Across every sector of society, people are facing unsettling questions about their worth and their place in a shifting world.

The rapid acceleration of AI comes with a profound wave of anxiety. Across every sector of society, people are facing unsettling questions about their worth and their place in a shifting world.
This session covers the foundational knowledge of GPUs, including their architecture, functionality, and applications in computing. It provides an introduction to GPU computing through the lens of the Marlowe SuperPod and prepares learners for advanced topics such as GPU-accelerated data science and machine learning.
Location: CoDa W401

This session covers the foundational knowledge of GPUs, including their architecture, functionality, and applications in computing. It provides an introduction to GPU computing through the lens of the Marlowe SuperPod and prepares learners for advanced topics such as GPU-accelerated data science and machine learning.
Location: CoDa W401
While current evaluation practices rely on static benchmarks, these methods face fundamental efficiency, reliability, and real-world relevance challenges. This talk presents a path toward a measurement framework that bridges established psychometric principles with modern AI evaluation needs. We demonstrate how techniques from Item Response Theory, amortized computation, and predictability analysis can substantially improve the rigor and efficiency of AI evaluation. Through case studies in safety assessment and capability measurement, we show how this approach can enable more reliable, scalable, and meaningful evaluation of AI systems. This work points toward a broader vision: evolving AI evaluation from a collection of benchmarks into a rigorous measurement science that can effectively guide research, deployment, and policy decisions.
