TD-MPC-Glass is a field report on trying to beat a strong model-based RL agent (TD-MPC2) at the architecture/algorithm level, under a strict fair protocol: single-variable, compute-matched, pre-registered peak-AND-final CI gates, and a mechanism-check before every multi-week campaign. The headline result is a careful negative one — explicit “abstraction” is largely redundant with what a trained self-predictive world model already learns — and the transferable contribution is the methodology that found that out cheaply.

These posts walk from the first JAX/Flax reimplementation through the full null parade and the value-probe that explained it. For the code and the full campaign record, see the links below.