On world models
What fascinates me about world models is the possibility of a system that understands space well enough to imagine what happens next.
What is a world model?
A world model is a representation of an environment. In its predictive form, it learns how that environment changes, including in response to actions. Here, “imagining” means predicting possible futures from what it has observed and what we might do. Ha and Schmidhuber's World Models is an early example of learning this from experience.
Imagine a chair in a hallway. From one photograph, we see only part of it. A useful model might predict the view from behind the chair, or what happens if we push it. Those are different demands, and a system that handles one may fail at the other.
The word “world” is quite generous here. The environment can be one game or one room. The model only needs to represent enough of it for the task.
Why I keep coming back to them
With ArBit, I wanted virtual objects to stay attached to physical places. Gaussian Splatting lets us move through a scene captured in photographs. World models take that curiosity further: how much of a world's behaviour can a model learn?
I want a model that can keep track of the chair after I look away, anticipate where it could move, and use that prediction to help decide what to do. Space, memory, and consequences become part of the same problem. That's a lot of what draws me to this field.
It also gives an agent somewhere to practise. Ha and Schmidhuber demonstrated this in a limited game setting: an agent learned inside a model-generated environment, then ran in the original game. Learning from an imagined consequence is a pretty interesting idea.
The hard part is making those predictions hold up. Genie 3 generates views in response to your controls, but its authors describe limits on interaction and memory. A convincing few seconds does not establish reliable physics.
I want to explore how far that understanding can go: whether a model of a place stays useful when the camera moves, time passes, or someone acts.
Update: what do we mean by understanding?
June 5, 2026.
Fei-Fei Li's A Functional Taxonomy of World Models, published in June, helped me be more precise about this. She separates three functions: rendering, simulation, and planning. My reading, using the same hallway:
| Function | Output | In the hallway |
|---|---|---|
| Renderer | A view | Show the chair from the doorway. |
| Simulator | A state | Predict where the chair moves after a push. |
| Planner | An action | Choose a path around the chair. |
Here, state means properties such as position, shape, and speed. These functions can be combined; in model-based reinforcement learning, a controller uses the world model to choose actions.
For me, this puts more weight on the word “understands” in the opening paragraph. Rendering a plausible hallway is one thing to get right. Predicting what a push does to the chair asks more of the model. Choosing a useful action also requires a goal.
The connection between those abilities is what interests me most. I want to see how much of the same learned knowledge can support all three.