Learning Latent Action World Models In The Wild

Published
Source
arXiv
Paper number
106
Field
World Models
arXiv ID
2601.05230

Key points

  • Most successful world models rely on explicit action labels, which limits scalability and generality when they are applied to massive amounts of unlabeled real-world video data.
  • Existing latent action models (LAMs) have been confined to limited environments, which has produced action spaces with weak transferability and poor generalization to diverse real-world scenarios.
  • Wild videos introduce problems such as poorly defined actions, environmental noise, and the lack of a shared embodiment, which makes it difficult to extract meaningful action representations.
  • Using the frozen V-JEPA 2-L visual encoder, the framework jointly learns a world model and inverse dynamics models from large-scale unlabeled in-the-wild video datasets.
  • It investigates information regularization methods, including sparsity, noise injection in a VAE-like form, and discretization, to control the content of continuous latent actions and prevent irrelevant information from being encoded.
  • It evaluates the learned latent action space using intrinsic metrics such as future leakage and action transferability, as well as practical utility on downstream planning tasks through a lightweight controller and the Cross-Entropy Method.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)