Evidence-Backed Video Question Answering

Published
Source
arXiv
Paper number
613
Field
Computer Vision
arXiv ID
2607.11862

Key points

  • E-VQA defines a new task that requires three outputs: an answer, temporal evidence, and spatial evidence in the form of masklets.
  • ST-Evidence is the first human-verified spatiotemporal grounding benchmark, with both generative and objective variants.
  • It shows that state-of-the-art models such as Qwen3-VL, OpenAI-o3, and Gemini-2.5-Pro have decoupled QA accuracy and grounding ability.
  • An automatically generated 160K-scale ST-Evidence-Instruct pipeline bridges the gap between reasoning and grounding.
  • Fine-tuning UniPixel 7B brings large gains, with t-mean up by 27.2 and J&F up by 13.8.
  • Scaling alone is not enough; the task requires fundamental advances in architecture, data, and training objectives.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)