Stanford
University
  • Stanford Home
  • Maps & Directions
  • Search Stanford
  • Emergency Info
  • Terms of Use
  • Privacy
  • Copyright
  • Trademarks
  • Non-Discrimination
  • Accessibility
© Stanford University.  Stanford, California 94305.
Why Governing World Models Is AI's Next Big Policy Challenge | Stanford HAI
Skip to content
  • About

    • About
    • People
    • Get Involved with HAI
    • Support HAI
    • Subscribe to Email
  • Research

    • Research
    • Fellowship Programs
    • Grants
    • Student Affinity Groups
    • Centers & Labs
    • Research Publications
    • Research Partners
  • Education

    • Education
    • Executive and Professional Education
    • Government and Policymakers
    • K-12
    • Stanford Students
  • Policy

    • Policy
    • Policy Publications
    • Policymaker Education
    • Student Opportunities
  • AI Index

    • AI Index
    • AI Index Report
    • Global Vibrancy Tool
    • People
  • News
  • Events
  • Industry
  • Centers & Labs

Stay Up To Date

Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.

Sign Up For Latest News

Navigate
  • About
  • Events
  • AI Glossary
  • Careers
  • Search
Participate
  • Get Involved
  • Support HAI
  • Contact Us
news

Why Governing World Models Is AI's Next Big Policy Challenge

Date
August 04, 2026
Topics
Spatial Intelligence
Regulation, Policy, Governance
Busy downtown rush-hour driving scene from the perspective of a self-driving car

As artificial intelligence moves beyond language into the physical world through "world models," Stanford researchers warn that policymakers face an even steeper governance challenge than with large language models—and the window to get ahead of the technology is closing fast.

If crafting effective regulations for large language models has proven difficult, governing the next wave of AI will be exponentially more complex. A new policy brief from the Stanford Institute for Human-Centered AI (HAI) marks the first comprehensive examination of how to govern “world models”—AI systems that don't just process language but build working representations of physical environments to predict how they change in response to action.

World models are already emerging from research labs into commercial applications, from crisis response systems to autonomous vehicles to robotic manufacturing—with profound implications for everything from privacy to national security.

In this conversation, HAI Associate Director and Hoover Senior Fellow Amy Zegart, HAI Founding Director Fei-Fei Li, and HAI Executive Director Russell Wald, three of the brief’s authors, discuss why world models demand urgent policy awareness and what makes governing them fundamentally different from current AI oversight efforts.

What exactly is a world model, and why should policymakers care about it now?

Fei-Fei Li: Spatial intelligence is the ability to perceive, reason about, and act in three-dimensional space, and world models are the foundation for achieving it. A world model is an AI system that builds an internal representation of an environment—like a city, a factory floor, or a disaster zone—and uses that representation to predict what happens when you take action in that environment. Unlike a language model that predicts the next word, a world model predicts physical consequences: what happens if you move this object, open that door, or reroute traffic around a collapsed bridge? World models will unlock true spatial intelligence, enabling AI to truly understand and navigate in the world around us. 

The reason policymakers need to pay attention now is that we're seeing rapid commercial development. Major tech companies on both sides of the Pacific—Google DeepMind, Nvidia, Tencent, and startups—are pouring resources into this space. These aren't just research prototypes anymore. They're heading toward real-world deployment in robotics, autonomous vehicles, infrastructure planning, and national security applications.

Russell Wald: And the governance gap is enormous. We're already struggling to regulate language models, which primarily deal with information risks—misinformation, bias, copyright violations. World models create physical risks. When an AI system's understanding of the physical world guides a robot, drives a vehicle, or informs emergency response decisions, errors don't just misinform—they can injure, destroy property, or cost lives. We need to get ahead of this technology before it's deeply embedded in critical systems.

"We need to get ahead of this technology before it's deeply embedded in critical systems," says HAI Executive Director Russell Wald. | SF Photo Agency

You outline three functional categories of world models: renderers, simulators, and planners. Why does that distinction matter for policy?

Li: Because the governance requirements are very different depending on what the system actually does. A renderer generates images or video—it makes a world appear realistic. That's valuable for architecture, design, virtual training. But appearance isn't the same as accuracy. A renderer might produce a beautiful image of a new hospital wing without capturing whether the structure is sound or how people would actually move through it.

A simulator goes deeper—it models the underlying physics, geometry, and dynamics. It needs to behave in structurally consistent ways because engineers will use it to test bridge designs under wind stress or train robots for warehouse work.

A planner goes furthest—it determines what action an agent should take. That's the capability that powers autonomous vehicles, disaster response robots, systems that must operate in unstructured, changing environments.

Wald: The policy implication is that we can't regulate world models as a single category. The closer the system gets to safety-critical decisions or physical action, the more rigorous the evaluation needs to be.

What makes world models different from regulating language models or even agentic AI?

Zegart: With language models, we're focused on content. Is it accurate? Biased? Harmful? With agentic AI, we focus on authority—which actions is the system permitted to take and under whose oversight?

World models add a third dimension: the validity of the simulated environment itself. When a world model serves as a stand-in for physical reality, a convincing but flawed environment can propagate errors across every system that relies on it.

Here's a concrete example: Imagine an autonomous vehicle that's trained in a simulated city and then tested using that same simulation to verify it's safe. If the simulation underestimates how slippery roads get in rain, the vehicle will learn to drive too fast in wet conditions and will still score well in testing because the test uses the same flawed model. The vehicle looks ready for deployment, but the real road will expose the error—potentially fatally.

"We can't shortcut real-world validation just because a system looks good in simulation," says HAI Founding Director Fei-Fei Li. | SF Photo Agency

You identify several major risks in world models—privacy erosion, liability challenges, market concentration. Which concerns you most?

Zegart: For me, it's the national security implications. World models are profoundly dual-use. The same system that navigates humanitarian aid through a disaster zone can guide a weapon through contested terrain. Modern security competition increasingly hinges on the ability to model environments, rehearse contingencies, and operate autonomous systems with limited human oversight.

What makes this especially challenging is that world models could lower the barrier to entry for capable autonomous military systems. Historically, military advantage has required expensive platforms and extensive operational experience. World models could allow less-resourced actors—including adversaries and non-state groups—to test weapons and tactics in synthetic environments, narrowing experience gaps and enabling cheaper systems at scale.

The technology is also likely to become a major arena of U.S.-China competition. China has made embodied intelligence a national priority in its latest five-year plan, directing state funding and shared data infrastructure toward physical AI. The U.S. has led largely through private labs and academic research, but Congress and the White House are now considering more coordinated approaches.

Li: I'm concerned about concentration and access. World models depend on a type of data that's fundamentally different from what trained language models—action-labeled interaction data. Robot trajectories, fleet logs, teleoperation records. This data can't be scraped from the internet. You have to generate it by operating physical machines in real-world environments.

That creates enormous barriers to entry and gives advantages to a small number of well-capitalized firms that can deploy systems at scale. Competitors, startups, universities, and public institutions can't catch up.

Wald: I agree. And this isn't just about market competition. If governments depend on proprietary simulations for training autonomous systems or planning infrastructure, but can't inspect or replace those systems, critical public capabilities become dependent on private infrastructure controlled by a few firms.

We need public investment to broaden access. Shared pools of robot trajectory data, teleoperation logs, and public-interest simulation environments should be explicit targets of initiatives like the National AI Research Resource led by the National Science Foundation.

NSF should coordinate shared infrastructure. Sector-focused agencies—Transportation, Energy, Health and Human Services—should contribute domain-specific testbeds. States can support real-world testing through procurement. 

"We have a chance to shape the conditions under which this technology develops before its trajectory locks in," says HAI Associate Director and Hoover Senior Fellow Amy Zegart. | David Gonzales

The brief notes that no existing benchmark gives policymakers an adequate basis to approve a world model for safety-critical deployment. How big a problem is that?

Li: It's a fundamental gap. Current benchmarks mostly measure visual quality, not whether a system understands physical dynamics. Newer benchmarks are starting to test physical reasoning—whether objects behave correctly, whether scenes stay consistent, whether skills learned in simulation transfer to real tasks.

But we're still in a research patchwork. And leading models now achieve near-perfect scores on simulated tasks even though their real-world reliability remains limited, making those benchmarks less useful for distinguishing among systems.

Until measurement science catches up, we must continue to require rigorous field testing. We can't shortcut real-world validation just because a system looks good in simulation.

What should policymakers do first?

Zegart: Three things. First, fund the measurement science. Direct NIST to develop evaluation methods for world models, with sector agencies defining relevant operating conditions. Without valid benchmarks, we're flying blind on safety.

Second, use procurement strategically. Government agencies are major customers for autonomous systems, robotics, and simulation infrastructure. Every procurement should require independent testing, access to evaluate the simulation environment, and real-world validation. Make those requirements standard across federal purchasing.

Third, address the data infrastructure gap. Make datasets and public-interest simulation environments explicit priorities for the National AI Research Resource.

Li: I'd add a fourth: Invest in cross-disciplinary research and training. World models sit at the intersection of computer vision, robotics, physics simulation, and domain expertise in everything from urban planning to crisis response to national security. We need researchers and policymakers who can work across those boundaries.

Wald: Stanford HAI's model of bringing together computer scientists, social scientists, legal scholars, policy experts, and domain specialists needs to be replicated across universities and government agencies. The governance challenges aren't purely technical, and the technical solutions aren't purely algorithmic. They require sustained collaboration across disciplines.

Your conclusion emphasizes that the choices made now will determine whether world models develop inside "a narrow commercial and security logic" or within a framework that serves public value. What's at stake in that choice?

Li: We’re at an inflection point: whether this technology strengthens economic resilience, expands scientific capacity, and supports safer embodied AI or whether it concentrates power, expands surveillance, and accelerates unsafe deployment in high-stakes settings.

The difference comes down to decisions being made right now: investing in shared infrastructure before concentration hardens, tying safeguards to deployment contexts, and building measurement science and public expertise.

Zegart: The window is narrow. Once proprietary systems become deeply embedded in critical infrastructure, commercial fleets, and military operations, changing course becomes so much harder.

But here's what gives me hope: Unlike with language models, where we're playing catch-up, we're identifying these governance challenges while world models are still emerging from labs into early deployment. We have a chance to shape the conditions under which this technology develops before its trajectory locks in.

"The World Model and Spatial Intelligence Era: Governing AI Beyond Language" was written by the following Stanford HAI faculty and staff: Daniel Zhang, Russell Wald, Ehsan Adeli, Elena Cryst, Daniel E. Ho, Caroline Meinhardt, Jiajun Wu, Amy Zegart, and Li Fei-Fei.

Share
Link copied to clipboard!
Authors
  • headshot
    Shana Lynch

Related News

The Complexities of Governing Mental Health AI
Caroline Yee, Caroline Meinhardt, Michelle Mello, Jane Paik Kim
Jul 24, 2026
News
digital face mental health illustration

Policymakers, academics, healthcare providers, AI developers, and patient advocates convened by Stanford HAI identify critical gaps in how we regulate AI tools used for therapy and emotional support.

News
digital face mental health illustration

The Complexities of Governing Mental Health AI

Caroline Yee, Caroline Meinhardt, Michelle Mello, Jane Paik Kim
HealthcarePrivacy, Safety, SecurityGenerative AIRegulation, Policy, GovernanceJul 24

Policymakers, academics, healthcare providers, AI developers, and patient advocates convened by Stanford HAI identify critical gaps in how we regulate AI tools used for therapy and emotional support.

How AI Is Helping States Cut Through Decades of Red Tape
RegLab staff
Jul 23, 2026
News

Stanford researchers scanned 500 million words of state law to reveal the stunning growth of reporting requirements, and are now giving governments a tool to clean them up.

News

How AI Is Helping States Cut Through Decades of Red Tape

RegLab staff
Government, Public AdministrationRegulation, Policy, GovernanceJul 23

Stanford researchers scanned 500 million words of state law to reveal the stunning growth of reporting requirements, and are now giving governments a tool to clean them up.

The AI Sovereignty Paradox: Should Countries Buy, Build, or Lease to Maintain Strategic Control of Their AI?
Shana Lynch
Jul 14, 2026
News

As nations invest billions to reduce reliance on foreign AI providers, a new Stanford HAI report surveys commercial sovereignty solutions and assesses the extent to which they meaningfully reduce dependencies on U.S. tech giants.

News

The AI Sovereignty Paradox: Should Countries Buy, Build, or Lease to Maintain Strategic Control of Their AI?

Shana Lynch
Government, Public AdministrationInternational Affairs, International Security, International DevelopmentRegulation, Policy, GovernanceJul 14

As nations invest billions to reduce reliance on foreign AI providers, a new Stanford HAI report surveys commercial sovereignty solutions and assesses the extent to which they meaningfully reduce dependencies on U.S. tech giants.