Skip to main content

When Does a Learned World Model Help — or Hurt? (cross-model result)

·301 words·2 mins

Short version: whether a learned world model helps a planner or hurts it is a property of the task, not of the world-model architecture. We now see the same tasks flip load-bearing vs redundant across two independent WM families — TD-MPC2 (latent-consistency) and Dreamer (reconstruction RSSM).

The result
#

We ablate the world model (strip its forward-dynamics learning) and measure the change in return, per task, in each framework:

taskTD-MPC2Dreamer
cheetah-runstripping the WM under planner-collection helps +45% (the “inversion”)vanilla beats stripped by +9.7% (WM helps)
walker-runnull — imposed structure is redundantvanilla ≈ stripped (null)

Cheetah is the task where the value function needs a large fraction of the latent (low compressibility); walker is value-sufficient (a tiny bottleneck already recovers most of the return). The world model matters on the former and is redundant on the latter — and this ordering is identical in both frameworks.

Why it matters
#

Most “does a world model help” debates are argued at the level of architectures. This says the more useful question is at the level of tasks: a cheap, checkpoint-time probe (a value-sufficiency bottleneck) tells you, per task, whether the learned WM has room to help — before you spend the compute. And on the tasks where the value head is already sufficient, the WM is not just neutral: under planner-collection it can inflate the planner’s own target variance (~3×) and actively destabilize learning.

What’s next
#

The diagnostic is prescriptive, not just descriptive: it points at a gated world model — one that measures its own value-sufficiency and rollout variance and down-weights itself on the tasks (and states) where it would otherwise hurt.

The detailed lab notebook lives on the project page; the underlying numbers and commit hashes are in the campaign’s public issues and ledger.