Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.
Sign Up For Latest News
Open to Stanford community members! We are hosting an interactive orientation featuring faculty insights, networking, and opportunities to explore AI and data science research, programs, and resources across Stanford.
Open to Stanford community members! We are hosting an interactive orientation featuring faculty insights, networking, and opportunities to explore AI and data science research, programs, and resources across Stanford.
As AI moves beyond language into systems that can perceive, understand, and act in the physical world, a new frontier is emerging: world models—AI systems that build and maintain working representations of real environments to predict how they change in response to action.

As AI moves beyond language into systems that can perceive, understand, and act in the physical world, a new frontier is emerging: world models—AI systems that build and maintain working representations of real environments to predict how they change in response to action.
Sessions run Wednesdays from 4:30–5:30 PM in CoDa E160. Each session features a different Stanford speaker; talk titles are announced by the organizers.

Sessions run Wednesdays from 4:30–5:30 PM in CoDa E160. Each session features a different Stanford speaker; talk titles are announced by the organizers.
HAI Weekly Seminar
The world we live in is inherently compositional: just like a sentence is built upon phrases and words, a visual scene comprises a collection of interacting objects and entities, which in turn are derived from the sum of their parts. This compositionality plays a critical role in our ability to understand the world, organize the acquired knowledge through a rich set of concepts, and easily adapt them to novel situations and environments. Essentially, it is considered one of the fundamental building blocks of human intelligence. How to incorporate such compositionality into AI models? How can we encourage neural networks to develop semantic understanding of our surroundings? And how can we leverage the emerging structured knowledge to improve in downstream tasks such as question answering or image generation? These are the questions that will be explored in the talk, in which I will present models for multi-step synthesis of and reasoning over multi-object scenes, describe their key design principles and underlying mechanisms, and illustrate the benefits they offer in terms of enhanced controllability, increased data-efficiency, and improved interpretability of their internal representations and reasoning process.
PhD Student in Computer Science, Stanford University
No tweets available.