Learning Latent Action World Models In The Wild
- Published
- Source
- arXiv
- Paper number
- 106
- Field
- World Models
- arXiv ID
- 2601.05230
Key points
- Most successful world models rely on explicit action labels, which limits scalability and generality when they are applied to massive amounts of unlabeled real-world video data.
- Existing latent action models (LAMs) have been confined to limited environments, which has produced action spaces with weak transferability and poor generalization to diverse real-world scenarios.
- Wild videos introduce problems such as poorly defined actions, environmental noise, and the lack of a shared embodiment, which makes it difficult to extract meaningful action representations.
- Using the frozen V-JEPA 2-L visual encoder, the framework jointly learns a world model and inverse dynamics models from large-scale unlabeled in-the-wild video datasets.
- It investigates information regularization methods, including sparsity, noise injection in a VAE-like form, and discretization, to control the content of continuous latent actions and prevent irrelevant information from being encoded.
- It evaluates the learned latent action space using intrinsic metrics such as future leakage and action transferability, as well as practical utility on downstream planning tasks through a lightweight controller and the Cross-Entropy Method.
Paper links
External research summaries. These are not HDATF publications or measured product results.