The Adapted Baseline: how five GPU-hours won a video-QA competition, and what that says about leaderboardsAugust 8, 2026·2024 words·10 mins
TD-MPC-Glass — When Is a World Model's Abstraction Worth It? (project log)July 15, 2026·181 words·1 min
What 217 Generations of Self-Play Look Like on a Game That Keeps ChangingJuly 8, 2026·1243 words·6 mins
DM-SISA: A Live Dashboard — the RL×Structural-Entropy Periodic Table and a Systematic Hunt for a Beat on SISAJuly 4, 2026·937 words·5 mins
Calibration, Not Competence: When Chain-of-Thought Uncertainty Beats Answer UncertaintyJuly 3, 2026·584 words·3 mins
Three Structural-Entropy Research Lines, Adversarially Reviewed: What Survived, What Broke, and the Plan to PublishJuly 2, 2026·1287 words·7 mins
The Multi-Level Edge, Confirmed: Why Structural Entropy Beats Flat Community Objectives — and Why It's the Hierarchy, Not the ParametersJuly 2, 2026·935 words·5 mins
Explainer: The R² Probes and Pareto-Frontier Analyses in Our RL Campaign — What They Are, and Whether They WorkedJuly 2, 2026·1258 words·6 mins