Evidence-Backed Video Question Answering
- Published
- Source
- arXiv
- Paper number
- 613
- Field
- Computer Vision
- arXiv ID
- 2607.11862
Key points
- E-VQA defines a new task that requires three outputs: an answer, temporal evidence, and spatial evidence in the form of masklets.
- ST-Evidence is the first human-verified spatiotemporal grounding benchmark, with both generative and objective variants.
- It shows that state-of-the-art models such as Qwen3-VL, OpenAI-o3, and Gemini-2.5-Pro have decoupled QA accuracy and grounding ability.
- An automatically generated 160K-scale ST-Evidence-Instruct pipeline bridges the gap between reasoning and grounding.
- Fine-tuning UniPixel 7B brings large gains, with t-mean up by 27.2 and J&F up by 13.8.
- Scaling alone is not enough; the task requires fundamental advances in architecture, data, and training objectives.
Paper links
External research summaries. These are not HDATF publications or measured product results.