Hy-Embodied-VLM-1.0: Efficient Physical-World Agents
- Published
- Source
- arXiv
- Paper number
- 620
- Field
- Computer Vision
- arXiv ID
- 2607.12894
Key points
- It organizes pretraining and post-training data around a three-stage capability taxonomy: understanding action-relevant states, reasoning about action transitions, and sequential and adaptive reasoning.
- It limits active parameters to 3B using the Hy3-A3B language backbone, the Hy-ViT2 visual encoder, and a mixture-of-experts architecture.
- It achieved the highest performance among similarly sized models on 19 of 38 benchmarks and improved by an average of 8.4% over the previous Hy-Embodied-0.5.
- By supporting multi-turn interaction and long-horizon reasoning with few active parameters, it can serve as a foundation model for latency-sensitive physical-world agents.
- The model with 3B active parameters only approaches the previous generation's model with 32B active parameters rather than surpassing it on every evaluation, and its top performance was limited to half of the 38 benchmarks.
Paper links
External research summaries. These are not HDATF publications or measured product results.