OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

Published
Source
arXiv
Paper number
985
Field
Computer Vision
arXiv ID
2608.21360

Key points

  • To reduce the problem of user behavior branching according to model responses, it predefined interaction paths from the original videos and evaluated models along those paths.
  • By reverse-engineering internet videos and investing at least 1000 hours of expert work, it constructed seven tasks, sixteen subtasks, and three real-world cases.
  • The best model, Gemini-3-Pro, scored only 66.4 out of 100, while the open model Qwen3-Omni-Instruct scored 51.2.
  • Understanding hand-gesture instructions, long-term memory, delaying responses until a target event, and maintaining goals across multiple conversation turns emerged as shared bottlenecks.
  • Because interaction paths were fixed in advance to enable static evaluation, it cannot reproduce every conversational path that branches as real users and models make free choices.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)