Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.
Sign Up For Latest News

This brief highlights the emergence of world models and outlines a first-of-its-kind governance and policy agenda for the technology.
This brief highlights the emergence of world models and outlines a first-of-its-kind governance and policy agenda for the technology.



This conversation spans the evolution of AI, exploring the opportunities for AI to augment human capabilities, particularly in healthcare, education, and scientific discovery. Fei-Fei discusses the limits of today’s AI, the importance of preserving human agency, and why society, not just the technology industry, needs to play a role in shaping its future.
This conversation spans the evolution of AI, exploring the opportunities for AI to augment human capabilities, particularly in healthcare, education, and scientific discovery. Fei-Fei discusses the limits of today’s AI, the importance of preserving human agency, and why society, not just the technology industry, needs to play a role in shaping its future.
Artificial intelligence (AI) tools for radiology are commonly unmonitored once deployed. The lack of real-time case-by-case assessments of AI prediction confidence requires users to independently distinguish between trustworthy and unreliable AI predictions, which increases cognitive burden, reduces productivity, and potentially leads to misdiagnoses. To address these challenges, we introduce Ensembled Monitoring Model (EMM), a framework inspired by clinical consensus practices using multiple expert reviews. Designed specifically for black-box commercial AI products, EMM operates independently without requiring access to internal AI components or intermediate outputs, while still providing robust confidence measurements. Using intracranial hemorrhage detection as our test case on a large, diverse dataset of 2919 studies, we demonstrate that EMM can successfully categorize confidence in the AI-generated prediction, suggest appropriate actions, and help physicians recognize low confidence scenarios, ultimately reducing cognitive burden. Importantly, we provide key technical considerations and best practices for successfully translating EMM into clinical settings.
Artificial intelligence (AI) tools for radiology are commonly unmonitored once deployed. The lack of real-time case-by-case assessments of AI prediction confidence requires users to independently distinguish between trustworthy and unreliable AI predictions, which increases cognitive burden, reduces productivity, and potentially leads to misdiagnoses. To address these challenges, we introduce Ensembled Monitoring Model (EMM), a framework inspired by clinical consensus practices using multiple expert reviews. Designed specifically for black-box commercial AI products, EMM operates independently without requiring access to internal AI components or intermediate outputs, while still providing robust confidence measurements. Using intracranial hemorrhage detection as our test case on a large, diverse dataset of 2919 studies, we demonstrate that EMM can successfully categorize confidence in the AI-generated prediction, suggest appropriate actions, and help physicians recognize low confidence scenarios, ultimately reducing cognitive burden. Importantly, we provide key technical considerations and best practices for successfully translating EMM into clinical settings.

This brief demonstrates how real-time monitoring can address critical gaps in the oversight of radiological AI tools.
This brief demonstrates how real-time monitoring can address critical gaps in the oversight of radiological AI tools.


As Chinese AI closes the capability gap, Washington and Silicon Valley debate open-weight models. Stanford HAI's James Landay says it's the right conversation framed the wrong way.
As Chinese AI closes the capability gap, Washington and Silicon Valley debate open-weight models. Stanford HAI's James Landay says it's the right conversation framed the wrong way.
