If a system can predict what happens after an action, it can compare choices before acting. World-model research asks how to represent environmental change and use predictions for learning and decisions.
A world model learns how an environment changes from observations and actions. If a robot pushes a cup, it may predict where the cup will move; an agent can compare the consequences of several actions. Predictions may concern images or compressed internal states. Their value depends on how they help a task.
Related routes learn useful representations for controllers, train policies on imagined trajectories, or search with a model before acting. Producing plausible videos is another capability. Keep three questions separate when reading: what is predicted, which component is trained, and where the final outcome is evaluated.
- Compress high-dimensional observations into representations useful for decisions.
- Generate training trajectories in a learned environment, reducing some real interaction needs.
- Predict consequences to support search and planning before acting.
Judge a world model by the decisions it supports and whether success inside the model survives evaluation in the actual task.
- 01Observe and encode
- 02Predict states
- 03Compare futures
- 04Act in the world
- 05Feedback and update
Research milestones
- 2018
Separate vision, memory, and control
World Models uses compressed observations and dynamics to support a small controller, including training inside a learned environment.
- 2020
Learn what planning needs
MuZero learns internal states for reward, value, and policy prediction, then uses tree search. It need not reconstruct future frames.
- 2025
Learn behavior on imagined trajectories
The formal DreamerV3 paper studies one configuration across task families. Follow how the world model, actor, and critic work together.
Key concepts
- Latent state
- A compressed internal representation of observations and history. Individual dimensions need not have readable meanings.
- Dynamics model
- A model predicting how the next state depends on the current state and action.
- Imagined rollout
- A trajectory generated by repeatedly predicting action outcomes from an internal state. Its feedback comes from the model.
- Planning / Policy
- Planning compares candidate action sequences at decision time; a policy maps observations or internal states to actions. They can be combined.
World Models and Dreamer
World Models combines compressed observations, dynamics prediction, and control. Follow its diagrams and demos, identifying every module's inputs and outputs.
Then inspect DreamerV3 and how the learned model participates in behavior learning. Relevant foundations include representation learning, sequence prediction, and RL policies and returns.
What to predict, and what to do with predictions
What is predicted? Images, compressed states, or other representations. Representation choices affect training difficulty and the information retained.
What is prediction for? Generating future frames, simulating interaction, training policies, and planning have different requirements. To judge whether actions improve, evaluate downstream task outcomes.
For image and video synthesis, see generation; for robot applications, see embodied AI.
Trace a world-model system
Draw the data and module relationships in World Models. Find a prediction failure and consider its effect on later decisions. Then try running or training under the project's conditions. Long training need not be your first encounter.
Use Awesome World Models by topic. Following one use case builds understanding more readily than treating every system called a world model as the same thing.
World-model papers and research notes
The world-model catalog includes classic projects and paper indexes. Purshow Notes also contains WAM and related technical notes to consult by question.
Representative papers
Begin with the World Models interaction diagram, then compare MuZero and Dreamer: they need different predictions and use them at different times.