Evaluation Infrastructure That Catches Its Own Lies: Lessons from a Mahjong AI CampaignJuly 2, 2026·1474 words·7 mins
Jumpy World Models from Scratch: γ-Models, TD-Flow, and Compositional Planning (vs Dreamer & TD-MPC) June 18, 2026·1413 words·7 mins
What Kind of Thing Is Structural Entropy? A Role Taxonomy, and the LLM TurnJune 11, 2026·438 words·3 mins
Tagging 81 Papers So the Survey Can't Lie: a Tidy-Data Corpus for Structural EntropyJune 7, 2026·371 words·2 mins
Six Mirages: What 14 Iterations of Trying to Beat TD-MPC2 with Abstraction Taught UsJune 7, 2026·2915 words·14 mins
selib: a Standardized Library for Structural Entropy (and Optimizers that Beat the Originals)June 6, 2026·711 words·4 mins
Parity Was a Mirage: How Fixing Our Yardstick Revealed the RL Had Been Working All AlongJune 5, 2026·1541 words·8 mins
A 24% Gain That Might Be a Mirage: Multi-Relational Channel Attention, and How We're Testing It HonestlyJune 5, 2026·1372 words·7 mins
Two Roads to a Forecasting Paper: a Clean SOTA Beat, an Honest Negative, and What Each Needs to PublishJune 4, 2026·1546 words·8 mins
Reaching for Rank 1: The Parity Trap, Distilling the Champion, and What the Mahjong Winners Actually DoJune 4, 2026·1243 words·6 mins