Core argument: Every world model runs on the same skeleton: look at the current state of an environment, take in an action, predict the state that follows. The four main technical routes differ in exactly one design question — what should the model actually predict? Compressed hidden states (Dreamer-style imagination training), abstract meaning instead of pixels (JEPA), the raw video frames (video simulators like Genie 3), or the 3D stage behind the images (scene models like Marble). Each choice trades fidelity for efficiency, and each already has a working system behind it.
This is the third chapter of an ongoing series on world models. Chapter 1 asked why language models hallucinate; chapter 2 explained what separates a world model from an LLM. This one opens the box: how does a machine actually learn to rehearse the world in its head?
1. The short answer
A world model learns one reflex: given what I see and what I am about to do, predict what I will see next. Do that reflex well, and something remarkable becomes possible — the model can chain predictions forward without acting in reality at all. Roll the reflex out a hundred steps and you get a simulated experience, an imagined trajectory. That is what researchers mean by imagination: not daydreaming, just prediction, repeated until it looks like a story.
Everything in this chapter is a variation on one question. When the model predicts “what comes next,” what exactly is the it? Pixels? Compressed states? Abstract meanings? A 3D scene? The four main answers define the four technical routes, and each route has already produced a real system you can point at.
2. The flight simulator in the cockpit
Picture how airlines train pilots. Nobody hands a student a Cessna and hopes for the best. The student spends hundreds of hours in a flight simulator, practicing engine failures and crosswind landings, before touching a real plane. The simulator is a model of the world built for one purpose: let the trainee act, and show consequences, cheaply.
A world model is an attempt to give an AI the same training ground. Practice the maneuver a thousand times inside the model, then perform it once in reality.
The analogy breaks exactly where the risk lives. A flight simulator’s physics are programmed by engineers and checked against real flight data for years. A learned world model discovers its physics from data, on its own. If the training data was thin, or the model’s capacity is too small, the internal simulator can be confidently wrong — the instrument panel lies, smoothly and consistently. A pilot trained on a lying simulator flies badly in ways that feel right. Much of the research in this field is, at bottom, the problem of auditing the simulator before trusting the pilot it trains.
3. The shared skeleton
Strip away the branding and every world model has the same moving parts. It receives a description of the current state — camera frames, sensor readings, a scene. It receives an action. It outputs a prediction of the next state. Trained on long recordings of the world, it gets better at this the way a chess player gets better at reading a board: not by memorizing every position, but by absorbing what tends to follow what.
Three abilities stack on top of that reflex. Prediction is the base. Planning comes next: because the model can roll the future forward, an agent can try five actions in its head and keep the one with the best imagined outcome. Imagination training comes last: once the model is good enough, it doubles as a practice hall, generating endless situations for an agent to learn from without touching reality.
The four routes below are four answers to the only design question that really matters: when the model predicts the next state, what should that state be made of?
Route 1: Predict compact states, and train inside them
The Dreamer line of research — now at version 4 — builds a world model that compresses what it sees into small internal codes, then predicts how those codes change when actions happen. In the paper Training Agents Inside of Scalable World Models (September 2025), Danijar Hafner and colleagues describe the current recipe: a tokenizer that squeezes video into discrete tokens, a dynamics model that predicts how the tokens evolve under actions, and a training technique called shortcut forcing that keeps the imagined rollouts consistent over long stretches. The agent then learns entirely inside this world — the paper reports that an agent trained this way reached the Minecraft diamonds benchmark, a task that demands a long chain of planned steps, and transferred to the real game without further tuning.
The appeal is efficiency. Internal codes are cheap to predict, so the model can imagine thousands of practice runs where a video approach would drown in pixels. The cost is opacity: the model’s “state” is a set of numbers no human can casually read.
Route 2: Predict meaning, skip the picture
Meta’s V-JEPA 2 takes the opposite bargain. Its name holds the clue: JEPA stands for joint-embedding predictive architecture, the research line Yann LeCun has championed for years. Instead of predicting pixels, the model predicts in an abstract representation space — a running summary of what is where and what is changing — on the theory that understanding means capturing exactly that, not how every texture is shaded. Skipping appearance frees the capacity for the parts that matter for action.
The shortcut has a headline result behind it: tuned for action with a small amount of robot footage, the model ran pick-and-place tasks on robot arms in labs it had never seen, with success rates around 65 to 80 percent and no robot-specific training. The previous chapter covered that experiment in detail; the point to keep here is what it says about the method. A model that never learned to draw a kitchen still learned enough about kitchens to work in one. And the line has kept moving: Meta released V-JEPA 2.1 in March 2026, pushing the same self-supervised recipe toward dense, patch-level features across images and video.
Route 3: Predict the video itself, and make it controllable
Google DeepMind’s Genie 3, previewed in August 2025 and since opened to the public through Project Genie, takes the most literal approach: predict the next frames, as video, in real time — 720p at 24 frames per second, interactive, generated live from a text or image prompt rather than pre-rendered. What makes it a world model rather than a video generator is consistency and control: environments hold together for several minutes at a time, and they have spatial memory — walk through a scene, turn around, walk back, and what you saw is still there instead of being re-rolled. You can act inside the prediction, which is the property video generators have historically lacked.
This route buys believability — the visuals are the best of the four — and pays for it in compute and in the physics problem from chapter 2: a model predicting pixels can paint a glass that shatters wrong.
Route 4: Rebuild the 3D stage behind the images
World Labs, Fei-Fei Li’s company, released Marble in November 2025 as the first commercial world model of its kind. Give it text or images, and it produces a navigable, persistent 3D world — represented as Gaussian splats, a technique for storing scenes as clouds of soft 3D points — that can be exported into other creation tools and game pipelines. Where Genie 3 generates the view, Marble generates the stage: geometry that keeps existing when you are not looking at it, and that other software can walk around in.
The trade runs the other way from route 3. Explicit 3D is consistent and portable by construction, but the visual richness and open-ended variety are harder to reach than in generated video.
In September 2026 the lab pushed the line further with Atlas, an “omni” world model: a single system that can generate camera-controlled video up to a minute at 1440p from one image, reconstruct full 3D worlds from a handful of shots, and steer a camera through what it has built. The stage-builder and the view-generator are starting to merge.
4. My take
Three readings, in descending order of confidence.
First: the four routes are best understood as four points on one dial — how much does the model throw away? Route 1 throws away almost everything and keeps dynamics; route 2 throws away appearance and keeps meaning; route 3 keeps appearance and pays in compute; route 4 keeps geometry and pays in variety. No choice is free. When people argue that one route “wins,” they are usually arguing about which discards matter least for their application.
Second, medium confidence: the end state is a stack, not a champion. The obvious hybrid is already visible in the research — an abstract planner (route 1 or 2) proposing actions, a video model (route 3) rendering the consequences, a 3D representation (route 4) guaranteeing the geometry stays put. World Labs’ Atlas, which folds generation, reconstruction, and simulation into one shipping product, is that convergence arriving early. The labs are converging on combinations because each layer covers another’s weakness.
Third: for anyone evaluating claims, the useful question is not “which route is it?” but “show me the simulator’s error bars.” Every route produces systems that are confidently wrong somewhere; the mature ones are the teams that can tell you where. That is the audit question I will keep applying as this series moves through the companies building these systems — which is exactly where the next chapter goes.
5. Questions people actually ask
Are these four routes competing standards? Less than they look. They share the same skeleton — state, action, next state — and differ in what the state is made of. Teams mix them: an abstract planner on top of a video model is already a common research pattern.
Which route is winning? Depends on the job. For cheap practice at scale, imagination training leads. For robots, latent prediction has the strongest results so far. For believable interactive worlds, video is ahead. For portable 3D content, scene models. The honest answer today is that the field is hedging, and the labs with money are hedging on purpose.
Is a world model just a game engine? The comparison is useful precisely because it fails. A game engine’s rules are written by engineers; a world model’s rules are learned from data. Learning is what lets it capture places no designer ever built — and also what makes it capable of being smoothly, confidently wrong. The engine is audited because a human wrote every line. The world model is an audit problem.
Sources
- D. Hafner et al., “Training Agents Inside of Scalable World Models” (Dreamer 4), arXiv:2509.24527, September 2025.
- Meta AI, “Introducing the V-JEPA 2 world model”, June 11, 2025; model paper arXiv:2506.09985.
- Google DeepMind, “Genie 3: A new frontier for world models”, August 2025.
- World Labs, “Marble: A Multimodal World Model”, November 12, 2025; “Atlas: A World Model for Spatial Intelligence”, September 2026.
Next in this series: who is actually building world models? A map of the labs and startups — from DeepMind and World Labs to the newest billion-dollar bets — and what each one is really betting on.