AI & Computingpreprint2026-08-18

From Implicit Latents to Explicit Physics: Retained Decoding with Budgeted Supervision in LeWorldModel

Open access0 citations

Abstract

Joint-Embedding Predictive Architectures (JEPA) train world models to imagine action consequences in latent space, but the physical quantities they encode remain implicit and unreadable from the outside. We extend LeWorldModel (LeWM) with a shared decoder head on its latent, retained at inference, that reads out 6-D block and end-effector positions. Naive joint training at natural supervision strength (one epoch) collapses the latent geometry: prediction loss reaches 20.3× (SIGReg) and 39.4× (TC-SIGReg) vs. a physics-free LeWM baseline at the same checkpoint. The physics objective competes with the anti-collapse regularizer. With 2-epoch training, budgeted physical supervision (λphys = 0.1, two-layer budget) limits the increase in prediction loss to +9.5% relative to the 2-epoch physics-free baseline, compared to +437% without the budget. The accuracy gain comes from reshaping the latent itself, not a stronger decoder. Under a fixed decoding protocol, block-position error drops 4.1×, and physical quantities become more linearly readable in PCA. SIGReg favors single-frame decoding accuracy, while TC-SIGReg favors multi-step rollout stability. Physics-CEM optimizes closed-loop planning directly on decoded block coordinates, linking JEPA world models to downstream control.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-18

Authors: Bangjun Huang