TD-MPC-Glass, Part 17: The +60% Was a Hover — Reward-Hacking on PandaPickCube, and What Actually Solves It
Part 12 reported the project’s one robust win over flat TD-MPC2: a jumpy (k-step macro) world model beating vanilla by +60% on PandaPickCube and +161% on Pick-Orient, CI-separated on peak and final, every seed. It was the positive result the whole campaign had been chasing. Then the user asked the one question we hadn’t: “can I see the video?” The video showed the gripper hovering over the cube, never closing, never lifting — for the entire episode. The “+60%” was real on the number and empty on the task. This is the post-mortem, and the fix.
1. The catch: a video, not a metric
The +60% was measured on episode return under a dense shaping reward. Returns rose, CIs separated,
seeds agreed — every box the protocol checks was ticked. But none of those boxes ask did the robot
pick up the cube? When we rendered 40 episodes of the trained policy (bare policy and the planned
MPPI/jumpy variants), the answer was unambiguous: 0% real picks. The cube was never lifted past
~7cm of its 20–40cm target; box_target (the real task signal) peaked at 0.19, never the 0.9 success
threshold. The gripper just hovered.
This is textbook reward-hacking: optimizing a proxy (dense return) to the detriment of the true objective (pick the cube), because the proxy and the objective came apart. (Important caveat, fully developed in §6: this was measured at 1–1.5M steps. A stock PPO at full budget — ~30M steps — actually solves the task at 66%, and is itself 0% until ~28M. So the hover is as much an under-training symptom as a reward flaw; the reward is solvable, just slow. Read on.)
2. How the dense return is actually computed (the scale that hid it)
The PandaPickCube per-step reward (mujoco_playground):
r = 4.0·gripper_box + 8.0·box_target·reached_box + 0.25·no_floor + 0.3·robot_qpos
Each base term is a 1 − tanh(·) shaping signal in [0,1]. So max ≈ 12.55/step. The training/eval
return is summed over a single continuous 1000-step episode: the evaluator wraps the env with brax
EpisodeWrapper(episode_length=1000), and the inner env has no step-counter termination, so done
stays 0 until step 1000 (verified — there is no auto-reset inside the rollout; the agent gets one
long grasp→lift→place→hold attempt). So the logged-return ceiling is ~12,550 (12.55 × 1000), and:
| what | logged return | box_target_max |
real success |
|---|---|---|---|
| vanilla TD-MPC2 @1.5M (3 seeds) | ~2,000–2,670 | 0.000 | 0% |
| theoretical max (1000-step) | ~12,550 | 0.9+ | — |
Vanilla plateaus at ~2,500 with box_target_max = 0.0000, i.e. it leaves ~10,000 reward on the
table by never picking. The gap-to-ceiling, with zero box_target, is the hover-hack — quantified.
And it retires an earlier worry that “2967 > the max”: the max over the 1000-step eval is ~12,550, not
1882, so both the vanilla 1854 and the jumpy 2967 are simply two hovers, one banking shaping reward
slightly faster than the other. The “+60%” is one hover out-hovering another.
3. Well-ORDERED, badly-SHAPED (the mechanism)
Is the reward just wrong? No — it’s well-ordered. Scripting the three regimes and measuring return (per 150-step episode here, for clarity):
| regime | return | box_target |
from gripper_box |
from box_target |
|---|---|---|---|---|
| hover (TD-MPC2) | 315 | 0.00 | 279 (89%) | 0 |
| grasp, no place | 674 | 0.48 | 461 | 174 |
| full success | 965 | 0.91 | 478 | 445 |
Over a 150-step episode, a real success earns 3.1× the hover’s return — at that horizon the
ordering is correct. The problem is shaping, and it gets worse with horizon: the hover banks
89% of its return from the dense gripper_box proximity term just by floating near the cube, while
the big box_target reward is gated behind reached_box (gripper must first close to <1.2cm) and
only pays after grasp+lift+place. So a gradient learner climbs the smooth, ungated proximity hill into a
hover local optimum. And at the actual 1000-step training/eval horizon the pathology inverts the
ordering entirely — §4a shows return becoming anti-correlated with success, because dwelling near the
cube for 1000 steps out-scores spending those steps actually moving it to the target. Mis-shaping, not
mis-ordering, drives the hack.
4. The task is solvable — a heuristic controller proves it
If gradient RL can’t escape the hover basin, can anything solve PandaPickCube? We took the
“learning beyond gradients” route: instead
of training a network, iteratively refine a programmatic controller from feedback (telemetry,
per-phase success, videos). The result (controller v9+):
- Real grasp → lift → place, video-verified, ~9% success (
box_target ≥ 0.9), 99% grasp, lift 0.69, over 256 envs × multiple seeds. - The real wall was cube orientation —
rot_errcappedbox_targeteven on good grasps; cracked with analytic level-gripper IK (compute a wrist pose that keeps the cube level), plus a long settled grasp-hold and a split lift→transport→place carry.
So PandaPickCube is solvable; gradient TD-MPC2 just reward-hacks it. Success videos:
HL_v9_SUCCESS_env6.mp4, env21.mp4.
4a. The heuristic learning curve — in RL units
A programmatic controller doesn’t “train,” so to compare it to TD-MPC2 we evaluate every controller
iteration under the exact 1000-step eval_pi protocol TD-MPC2 is scored on (one continuous 1000-step
rollout, dense reward summed, n=64 episodes). That puts the heuristic’s “return” on the same axis:
| iter | version | return (1000-step) | real success | grasped |
|---|---|---|---|---|
| 1 | v1 — hover-grasp, no level-IK | 4403 | 0% | 0.89 |
| 4 | v4 | 4488 | 0% | 0.84 |
| 7 | v7 — dead-end 6DOF orient | 4549 | 0% | 0.84 |
| 8 | v8 — analytic level-IK (cracks the 0% wall) | 3069 | 1.6% | 0.81 |
| 9 | v9 | 2880 | 6.3% | 0.81 |
| 10 | v9_rhold (BEST) | 3406 | 6.3% | 0.83 |
| 11 | v11 — xy-servo | 3522 | 3.1% | 0.97 |
| — | TD-MPC2 vanilla @1.5M | ~2500 | 0% | ~0 |

The striking finding: at the eval/training horizon, return is anti-correlated with real success.
The highest-return controllers (4400–4550) score 0% success — they grasp the cube and hold it low,
banking the ungated gripper_box term plus robot_target_qpos (arm near home) for all 1000 steps
without ever lifting it to the elevated target. When the controller learns to actually lift-and-place
(v8→v10), return drops to ~2,900–3,400 — moving the cube to target sacrifices that steady dwell
reward — but it earns real success. TD-MPC2’s ~2,500 hover sits below even the real-success HL
versions, at 0% success. So in this reward, maximizing 1000-step return actively works against the
task — the cleanest possible picture of why gradient RL reward-hacks here. The HL line’s value is not
a number above 2,500; it’s that it converts cheap dwell-return into actual picks. Curve:
demo_videos/hl_learning_curve.{png,csv} + HL_CURVE_README.md.
4b. Pushing success higher: the wall is cube orientation, not grasp precision
Continued iteration set an honest ceiling: best = 0.094 real success (v9_rhold, 4-seed mean at
n=256; peak ~0.10). The instructive part is what didn’t work: iter-11 broke the long-standing
grasp-precision cap (reached_box 0.78 → 0.97 via a live-xy descent servo) — yet success fell,
proving reached_box was not the binding constraint. rot_err is: the cube rotates between the
parallel fingers during close/lift (~7–10° tilt), capping box_target below 0.9. An exact level-frame
target (iter-12) tied baseline — confirming the tilt is grasp dynamics, not the commanded pose. The
remaining untried lever (logged for resume): a physical regrasp/de-tilt settle on the table before lift.
5. What we’re doing about it: reward engineering to actually solve it (live)
The diagnosis (§3) names the fix: break the hover local optimum so gradient RL can reach the cube. We
are sweeping reward re-engineerings — train on the engineered reward, but score success on the
original box_target ≥ 0.9 task (a fair test) — including:
- Down-weight proximity (
gripper_box4→1) so hovering pays far less. - Ungated lift bonus — reward cube height above the table directly, not gated behind
reached_box, so “pick it up” has gradient before a perfect grasp. - Grasp bonus — a discrete reward the instant
reached_boxlatches, to bridge the gated cliff. - Potential-based staging — reach → grasp → lift → place as a shaped potential (policy-invariant, so it can’t introduce new hacks).
The bar: real success rises above 0 for vanilla and/or jumpy TD-MPC2, on the true metric.
Results (vanilla TD-MPC2, seed 1, eval/50k; all numbers read from *_realsuccess.csv):
Grasp rate is shown for both the bare policy and the planner (MPPI), since the planner can stumble into a grasp by search even when the policy can’t reproduce it:
| variant | config (weights) | grasp: policy / planner | peak box_target |
success |
|---|---|---|---|---|
| baseline | original reward | ~0 / ~0 (flickers @1.4M+) | 0.085 | 0 |
| V1 | reweight only: gripper_box 4→1, bt 12 | 0.00 / 0.00 | 0.00 | 0 |
| V2 | + ungated lift=4 (gripper_box=1) | 0.60 / 0.33 @150k → faded | 0.056 | 0 |
| V3 | V2 + grasp bonus 2 | 0.00 / 0.00 | 0.00 | 0 |
| V4 | lift=8, gripper_box=2 (full 1M) | 0.40 / 0.33 → 0 by 1M | 0.014 | 0 |
| V5 | lift=10, grasp=4, gripper_box=0.5 | 0.00 / 0.33 | 0.028 | 0 |
| V6 | lift=6, grasp=6, gripper_box=0.5 | 0.00 / 0.00 | 0.00 | 0 |
Verdict: reward engineering broke the hover trap but did not make TD-MPC2 solve the task. No variant reached real success > 0. But the sweep is informative:
- The ungated lift bonus is the one lever that works behaviorally. With a moderate proximity term
kept (V2:
gripper_box=1, ungatedlift=4), TD-MPC2 starts grasping in 60% of eval episodes by 150k — versus the original reward, where it never grasps before ~1.4M. Returns also dropped from ~2,500 (hover) to ~300–650, confirming the hover gravy-train was removed. The gating, not the ordering, is the trap — confirmed. - But grasping never consolidates. It’s intermittent and fades as the policy converges (V2 →
0.20, V4 → 0), never completing a lift→place, so
box_targetnever approaches 0.9. - Two dead ends, both informative. Cutting
gripper_boxto 0.5 (V5/V6) leaves the bare policy never grasping (the planner brushes it — V5mppi_reached=0.33 — but the policy can’t reproduce it), so the reach gradient is needed to get near the cube. And a discrete grasp bonus hurts (V3 → 0 grasps): it becomes its own little hack target rather than a bridge.
So reward shaping can knock the policy out of the hover basin, but it can’t manufacture a reliable grasp — and grasp reliability is the binding wall. That lines up exactly with the heuristic track (§4b): even a hand-coded controller with analytic level-IK tops out at ~9% because of grasp precision and cube-orientation dynamics. The cube is small and the grasp is unstable; that’s a dynamics/precision problem, not a reward-shape problem.
Final consolidation test (V7/V8, to 1M): stronger lift also fails. Adding more lift weight on the
good V2 base did not make the grasp stick — V8 (lift=12) peaked pi_reached=0.40 then faded to 0 by
900k (the same transient pattern), and V7 (lift=8, bt=15) never grasped at all. So across 8
reward variants up to 1M steps, real success stayed exactly 0 — the no-solve verdict is final, and the
bottleneck is grasp stability, not exploration or reward shape. One last probe: the jumpy
(temporal-abstraction) world model on the V2 reward config — if planning over macro-steps can’t
consolidate the grasp either, that closes the question for this task. It shows the exact same
pattern — over the full 1M steps, jumpy_reached intermittently brushes 0.33–0.67 via macro-step
planning but never sustains across consecutive evals, box_target ≤ 0.012, final and peak success
= 0 — i.e. temporal abstraction does not cross the grasp-stability wall either. So every
lever we have (reward shaping ×8, temporal abstraction) breaks the hover trap but none manufactures a
reliable grasp on this small-cube task; the honest conclusion is that PandaPickCube’s wall is grasp
dynamics, not reward design or planning horizon — which is precisely why the hand-coded controller
(§4) needed analytic level-IK to reach even ~9%. (Jumpy run complete at 1M: success 0, same
intermittent-grasp-never-consolidates pattern as every other variant. Reward-eng track closed.)

Honesty note: the box_target success metric was left untouched (raw + gated) throughout, so success
is measured on the true task; only the training reward was engineered. All grasp numbers are read from
the *_realsuccess.csv sidecars — both the policy (pi_reached) and planner (mppi_reached) columns,
distinguished above (an earlier draft conflated the two for V5). Curves:
demo_videos/reward_eng_curves.png, RESULTS.json.
Revised by §6 (important): the official PPO baseline below shows grasp reliability is achievable (99.6%) and the task is solvable (~66%) at ~18× this budget. So the §5 reading — “the wall is grasp dynamics, not reward/budget” — was too strong. The reward-engineering null is real at 1M steps, but the true limiter at that budget is budget + algorithm, not an unfixable reward or an intrinsic grasp-dynamics wall. Read §5 as “no reward tweak rescued TD-MPC2 at 1M,” not “the task is unsolvable.”
6. The decisive baseline: official PPO solves it — at ~18× the budget
Everything above used TD-MPC2 at 1–1.5M steps. The missing control: can the official
mujoco_playground recipe solve PandaPickCube at all? We ran stock brax PPO (its tuned PandaPickCube
config — 2048 envs, target 20M steps, policy 32×4) on the unmodified reward, scored with the same
box_target ≥ 0.9 metric (n=256, 1000-step eval). It does — but late, and only at full budget:
| env steps | success | grasp (reached) |
mean max box_target |
|---|---|---|---|
| 8.2M | 0.00 | 0.24 | 0.013 |
| 13.1M | 0.00 | 0.85 | 0.073 |
| 18.0M | 0.00 | 0.996 | 0.105 |
| 27.9M | 0.00 | 0.996 | 0.250 |
| 29.5M | 0.074 | 0.996 | 0.638 |
| 31.1M | 0.527 | 0.996 | 0.852 |
| 32.8M (final) | 0.660 | 0.996 | 0.887 |
(brax overshoots the 20M target to ~33M in its step accounting.) PPO reaches 66% real success, and the trajectory is telling: it learns reach → grasp (99.6% by 18M) → place (only ~29M+), grinding through the very same grasp-but-don’t-place plateau that traps low-budget agents.
Three honest corrections to this post:
- The reward is not “broken.” Standard tuned PPO climbs the dense reward all the way to real picks. The shaping is suboptimal (it parks the policy in a no-place plateau for ~28M steps) but solvable.
- TD-MPC2/jumpy’s 0% was mostly a budget failure. PPO is also 0% real success until ~28M — >18× the 1.5M we gave TD-MPC2. At 1.5M nothing solves this task; what we read as a “reward-hack” is largely “nowhere near converged.” (TD-MPC2 is usually far more sample-efficient per step, so it’s not settled it couldn’t solve this at a matched ~30M budget — we never ran it that far.)
- Grasp reliability is not an intrinsic wall. PPO grasps 99.6%; the heuristic’s ~0.78-grasp / 9%-success ceiling (§4) was an open-loop-controller limitation (it can’t react to the cube tilting), not a property of the task.
What still stands: the “+60% jumpy” headline was meaningless (both 0% real success at 500k); the metric-vs-task lesson holds (return looked great while the robot did nothing); and reward engineering at 1M genuinely didn’t rescue TD-MPC2 — but the reason is budget, not an unfixable reward.
Literature note (provenance of the “100%” claim). The MuJoCo Playground paper’s headline “robustly
grasps and lifts, 100% in real-world trials” is for the vision / Cartesian pick-cube variant (Madrona
rendering + domain randomization + a stochastic gripping delay) — not the state-based PandaPickCube
we use. For the state task the official recipe tracks only dense episode_reward (the manipulation
notebook plots no strict success rate or box_target threshold). That is precisely why the
reward-hacking went unnoticed upstream too: with no task-true success metric, a high-return hover looks
like progress. Our run is, as far as we can tell, the first to score state-PandaPickCube PPO on a
strict box_target ≥ 0.9 metric — and it does solve it (66%), just much later than the dense reward
would suggest.

SAC (unofficial config) gets stuck where PPO escaped. Our SAC reached grasp 99.6% but 0% real
success through 20M — box_target capped at 0.276, i.e. it learns to grasp but never places, never
escaping the grasp-but-don’t-place plateau that PPO broke out of at ~29M. (Caveat: this is not an
official tuned manipulation SAC config — mujoco_playground ships PPO for manipulation — so read it as
“a standard off-policy baseline at this budget,” not a definitive SAC verdict.) A longer PPO run (80M-target, ~113M
brax steps, 24 evals) answers whether PPO reaches 100%: it does not — it plateaus at ~80–83%.
Success crosses the threshold at ~20M (48%), reaches ~82% by ~25M, then stays flat at 80–83% for the
next ~85M steps (peak 83.2% @108M, grasp 100%, mean max box_target 0.948; final dips to 71.9%
@113M). So the longer run both crosses earlier and ceilings higher than the 20M-config run (66%@33M),
but the residual ~17% is a hard-instance / placement-precision tail, not a budget shortfall — even
a tuned on-policy PPO doesn’t fully solve every cube pose. Useful calibration for what “solved” means
here: ~83% is the practical ceiling for this reward+observation at scale.
Neutral offline dataset for InFOM (built). Since the original goal is to run InFOM/CompPlan on Panda
and TD-MPC2 rollouts would be unfair, we generated the dataset from the PPO policy (a standard
baseline): panda_pickcube.npz — 76,800 transitions, 512 episodes (256 expert + 256 exploratory),
75% success, OGBench layout. It ingests cleanly through ogbench.utils.load_dataset + InFOM’s
Dataset.create/.sample (smoke-tested); a full InFOM run just needs a small panda-pickcube branch in
env_utils.make_env_and_datasets to wire the eval env. So the InFOM-on-Panda pipeline is unblocked on
the data side.
6a. Update — InFOM does learn Panda offline, and a sister task (OpenCabinet)
InFOM-on-Panda (run complete, n=50): we wired the panda-pickcube branch and trained InFOM on the
neutral PPO dataset (500k pretrain + 250k finetune). Real success peaks at 0.74 @600k, then settles
to ~0.62–0.66 (final 0.62 @750k — a slight late-finetuning decline), with box_target ~0.75–0.82 and
reached ~0.82–0.92 throughout. Honest read: this is offline RL on a 75%-success expert dataset, so
InFOM ≈ recovers the demonstrator’s competence (~0.75 data → ~0.74 peak policy) — exactly what good
offline flow-occupancy RL should do. The point stands: offline RL solves Panda from a neutral, non-self
dataset — the original InFOM/CompPlan-on-Panda goal, achieved.
PandaOpenCabinet (sister task, identical reward shape). OpenCabinet uses the same reward
(4·gripper_box + 8·box_target·reached_box + …) with the “box” = cabinet handle pulled to a target
— a reach→grasp→pull skill. Official PPO solves it cleanly: 98.0% success (box_target ≥ 0.9,
n=256): 0% until ~13M → 84% @16M → 97% @26M → plateau 98.05% (peak @33M, final @75M), reached=0.98
— higher than PickCube’s 83% (no cube-orientation term to cap it) and success-onset earlier
(~13–16M vs ~29M).
But the HL loop fails OpenCabinet (0% success) — where it solved PickCube (9%). A
heuristic-learning controller (reach → grasp the vertical bar → pull along its slide-joint to target →
hold) reached grasp_ok 95% but reached_box only 7.4%, success 0%, across 12+ diagnosed iterations.
The binding wall is closed-loop: the handle sits at the Panda’s forward-reach frontier, where (1)
open-loop static-IK stalls ~10cm short (joint-5 saturates under position control — a joint-space
overshoot lever lifted grasp 28%→95% but couldn’t close the reach gap), and (2) the grasp site drifts
~1.5–2cm off the bar center during the pull, never holding the <1.2cm reached_box threshold and the
target simultaneously. PPO succeeds precisely because it learns closed-loop frontier reaching +
centered contact that an open-loop heuristic cannot reproduce. This is the cleanest demonstration of
the HL loop’s honest ceiling: it cracked PickCube’s grasp-orientation wall (open-loop-fixable → 9%)
but cannot cross OpenCabinet’s closed-loop-control wall (→ 0%). HL is a fast solver and diagnostic;
crossing a closed-loop wall needs learned control. (Details: hl_opencabinet/KNOWLEDGE.md,
demo_videos/OPENCABINET.md.)
6b. TD-MPC2 at scale on PickCube — grasps like PPO, but never places (0% to 24.3M)
We ran TD-MPC2 itself far past the 1.5M “reward-hack” budget — 6 seeds (3× vanilla latent-512, 3× small
latent-256) on free 3060s, the furthest reaching 24.3M steps. Read from the *_realsuccess.csv:
- Vanilla (latent-512) consolidates grasp like PPO:
reachedclimbs to 1.0 by ~13–15M (vs PPO’s 0.85@13M / 0.996@16M) — grasp-onset is comparable. The earlier “0% reward-hack” was a budget artifact. - But it never places.
box_targetstays flat at 0.04–0.07 from 5M all the way to 24.3M (best across all 6 seeds = 0.067; success threshold is 0.9) — 0% real success in every run. This is not “ran out of budget”: PPO’sbox_targetwas already 0.10 and climbing by 16M before it jumped to 83% at ~29M, whereas TD-MPC2’s shows no upward trend at all through 24.3M. It is genuinely stuck in the grasp-but-don’t-place plateau, with no convergence signal toward success near PPO’s onset budget. - Small (latent-256) lags further:
reachedstalls 0.4–0.6, never consolidates the grasp.
So on PickCube TD-MPC2 does not solve the task at any budget we reached (0% to 24.3M, flat box_target),
even though it grasps as early as PPO. The earlier “inconclusive / would-it-converge-at-29M” hedge is
retired: the plateau is flat with no trend, so the honest read is TD-MPC2 plateaus and does not convert,
not “we just didn’t run long enough.” PickCube is still a poor arbiter of sample-efficiency (dense
reward, very late success-onset); the clean head-to-head is on DMC — Part 18.
7. The lesson
- A metric is not a task. Every statistical box (CI, seeds, peak+final) was green while the robot
did nothing. Only a behavioral check (the video) and a task-true metric (
box_target ≥ 0.9) caught it. We’ve now wired real-success logging into the eval so this can’t hide again. - But “reward-hacking” can be a budget artifact in disguise. The most important correction in this post: a stock PPO at full budget (§6) solves the task, and is itself 0% until ~28M — so what looked like a fundamentally-hacked reward at 1–1.5M was largely under-training. Always run a strong, full-budget baseline before concluding “the reward/task is broken.” We almost didn’t.
- Dense shaping + a gated sparse jackpot = a long plateau, not a dead end. The ungated proximity term parks low-budget policies in a grasp-but-don’t-place plateau worth most of the return; PPO escapes it given ~30M steps. Suboptimal shaping, not a fatal one.
- “Learning beyond gradients” still earned its keep as a fast solver — a programmatic policy got real picks in hours of iteration vs PPO’s 33M-step grind — but its ~9% ceiling is an open-loop limit, not the task’s: closed-loop PPO reaches 66%.
Pointers
Correction banners on Parts 12 & 16. Reward source: mujoco_playground .../franka_emika_panda/pick.py.
Real-success logging + campaign: scripts/run_benchmark.py (_is_pickcube sidecar),
scripts/realsucc_campaign.sh. Heuristic controller + learning curve: hl_pickcube/ (controller.py,
KNOWLEDGE.md, LOG.jsonl), demo_videos/. Return-vs-success: demo_videos/goalB_results.json.