Stanford
University
  • Stanford Home
  • Maps & Directions
  • Search Stanford
  • Emergency Info
  • Terms of Use
  • Privacy
  • Copyright
  • Trademarks
  • Non-Discrimination
  • Accessibility
© Stanford University.  Stanford, California 94305.
Understanding Speech—Moment by Moment | Stanford HAI

Stay Up To Date

Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.

Sign Up For Latest News

Skip to content
  • About

    • About
    • People
    • Get Involved with HAI
    • Support HAI
    • Subscribe to Email
  • Research

    • Research
    • Marlowe (opens in new tab)
    • Research Programs
    • Grants
    • Student Affinity Groups
    • Centers & Labs
    • Research Publications
    • Research Partners
  • Education

    • Education
    • Executive and Professional Education
    • Government and Policymakers
    • K-12
    • Stanford Students
  • Policy

    • Policy
    • Policy Publications
    • Policymaker Education
    • Student Opportunities
  • AI Index

    • AI Index
    • AI Index Report
    • Global Vibrancy Tool
    • People
  • News
  • Events
  • Industry
  • Centers & Labs
Navigate
  • About
  • Events
  • AI Glossary
  • Careers
  • Search
Participate
  • Get Involved
  • Support HAI
  • Contact Us
news

Understanding Speech—Moment by Moment

Date
April 02, 2026
Topics
Sciences (Social, Health, Biological, Physical)

By Irmak Ergin, Jill Kries, and Laura Gwilliams

Understanding what someone is saying usually feels easy and automatic. But, we have all had the experience of misunderstanding someone, or failing to derive understanding entirely. Typically, research investigating speech comprehension does so by applying a “posthoc” measure of comprehension- that is, after the comprehension has happened, such as multiple-choice questions, rating scales, or summaries. These methods can capture overall understanding, but they miss the dynamic changes in comprehension as speech unfolds.

In our new paper, Measuring naturalistic speech comprehension in real time, we introduce a method designed to address this gap. We built a custom slider device that participants can use while listening to continuous, naturalistic speech, allowing them to report how well they understand what they are hearing in real time. The slider synchronizes with experimental software and provides millisecond-level readout, making it possible to generate a continuous behavioral trace of comprehension rather than a single score at the end.

To test whether this new measure works as intended, we evaluated it across three experiments. We asked the question: Does this continuous measure track comprehension at least as well as established post hoc methods, and perhaps better? Overall, the answer was yes. Across the study, slider responses captured fluctuations in understanding driven by manipulations such as speech rate (how fast the speech is) and information load (how surprising the content is), and the method was validated against existing measures.

At the same time, the paper highlights why standard post hoc measures are often limited. They rely heavily on memory, meaning they reflect not only comprehension, but also how much information a listener can retain afterward. Multiple-choice questions can be shaped by guessing or by the wording of the questions themselves. Summary-based measures introduce yet another complication, because they depend on both comprehension and the ability to reconstruct or retell what was heard.

One of the most promising aspects of the new method is its potential for cognitive neuroscience. The study shows that using the slider does not disrupt comprehension, making it well-suited for co-registration with neuroimaging methods such as electroencephalography (EEG) and magnetoencephalography (MEG), where we can look at the time-resolved neural dynamics of speech comprehension. This opens up an exciting new opportunity for language research. For the first time, it becomes possible to align the time course of the input speech signal with the listener’s changing experience of comprehension. It creates a path toward linking those behavioral dynamics to neural activity. Instead of asking only whether a listener understood something overall, we can begin to ask when comprehension succeeds, when it breaks down, and how those fluctuations relate to ongoing brain responses.

Although we apply the method here to speech comprehension, the broader idea extends well beyond language. If cognitive processes unfold over time, our measurements should too. This measure can enable researchers to study dynamic cognition in ways that static end-of-task measures simply cannot.

Read the paper!


This article is a part of the Stanford Data Science legacy publication. Read more about the HAI and Stanford Data Science merger.

Share
Link copied to clipboard!

Related News

Stanford Scientists Build an AI Lab Partner
Nikki Goth Itoi
Jul 09, 2026
News
DNA molecule spiral. 3d rendering

Biomni can analyze mountains of medical data, spot patterns humans might miss, and even design experiments—helping researchers make discoveries faster in the race to cure disease.

News
DNA molecule spiral. 3d rendering

Stanford Scientists Build an AI Lab Partner

Nikki Goth Itoi
Sciences (Social, Health, Biological, Physical)Jul 09

Biomni can analyze mountains of medical data, spot patterns humans might miss, and even design experiments—helping researchers make discoveries faster in the race to cure disease.

How AI Is Accelerating Scientific Discovery
Nikki Goth Itoi
Jul 08, 2026
News

New AI tools generate hypotheses, design experiments, and find patterns in data—transforming how scientists make discoveries across every field.

News

How AI Is Accelerating Scientific Discovery

Nikki Goth Itoi
Sciences (Social, Health, Biological, Physical)Jul 08

New AI tools generate hypotheses, design experiments, and find patterns in data—transforming how scientists make discoveries across every field.

Today's AI Talks Like “Nobody.” New Research Gives It Real Personality.
Jun 08, 2026
News
3D illustration of mirrored human profiles in blue and yellow layers

PsychAdapter lets researchers dial in on personality traits, age, and mental health characteristics to generate text that sounds like real individuals, opening the door to training simulations and personalized content.

News
3D illustration of mirrored human profiles in blue and yellow layers

Today's AI Talks Like “Nobody.” New Research Gives It Real Personality.

HealthcareGenerative AISciences (Social, Health, Biological, Physical)Jun 08

PsychAdapter lets researchers dial in on personality traits, age, and mental health characteristics to generate text that sounds like real individuals, opening the door to training simulations and personalized content.