Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
- Published
- Source
- arXiv
- Paper number
- 820
- Field
- Computer Vision
- arXiv ID
- 2608.02711
Key points
- It unifies 3D understanding, text-to-3D generation, instruction-based 3D editing, and text-based part generation in one architecture.
- It builds one of the largest multimodal 3D training corpora, with 87 million examples in total, including 25 million for understanding, 50 million for generation, and 12 million for editing.
- The Nano3D-v2 agent-based pipeline mass-produces geometrically consistent 3D editing data.
- The combination of Hunyuan3D-VLM for understanding and DiT for generation achieves state of the art on both generation and editing benchmarks.
- Experiments show cross-task synergy, where better generation and understanding improve editing quality.
Paper links
External research summaries. These are not HDATF publications or measured product results.