Incentivizing Vision Language Models to Search for Long Video Question Answering

Harsh Goel†, S P Sharan†, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep Chinchali
The University of Texas at Austin, USA
†Contributed equally to this work


ECCV Poster
Our ECCV Poster

Abstract

We introduce VSeek, an agentic framework that transforms long-video Question Answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek utilizes a natural language-driven search to identify relevant context within long videos and is post-trained with Reinforcement Learning (RL) to jointly formulate targeted search queries and reason over retrieved clips for LVQA. While RL post-training has revolutionized reasoning in symbolic domains such as mathematics and code, its application to long-video understanding remains hindered by a lack of verified rewards. To ensure that the retrieved context is relevant, we propose a novel neuro-symbolic approach that bridges open-ended natural language with discrete visual verification. Specifically, complex user queries are compiled into formal temporal logic specifications for systematically decomposing natural language questions into a definitive checklist of required atomic visual primitives, such as key objects and activities, along with their temporal ordering. These systematically derived grounding events provide the critical feedback signal for RL post-training, enabling dense, verifiable rewards based on the successful retrieval of these specific visual elements rather than relying entirely on outcome-only answer accuracy. By explicitly optimizing for this verifiable evidence-seeking behavior, VSeek improves Pass@1 scores by up to 8% and Pass@4 scores by 15% on long-video understanding benchmarks compared to base models.

Methodology


VSeek casts long-video question answering as a sequential decision process: a single VLM alternates between reasoning and retrieval, issuing natural-language searches over an indexed video until it has gathered enough evidence to answer. To post-train this behavior without relying on outcome-only accuracy, we introduce VETL, which compiles a question into a temporal logic specification and turns the visual primitives it requires into a dense, verifiable reward.



VETL reward computation



Key Capabilities


Verified Rewards Beat Exact Match

On three in-domain benchmarks, VSeek-VETL beats the base model it is post-trained from, the agentic baselines, and VSeek-EM, which sees only an exact-match reward. A 4B model trained this way overtakes GPT-5.2 on LongVideoBench and MLVU.

Gains Grow with Video Length

Dense, verifiable feedback matters most where uniform sampling breaks down. On the longest videos, VSeek-VETL gains 11.0 points over the base model on both LongVideoBench and MLVU, pulling clearly ahead of exact-match training.

Majority@4 accuracy on in-domain benchmarks
Accuracy by video length

Better Answers From Half the Frames

Average frames sampled per query. Where the static baselines always consume a fixed 64 frames and VideoTree scales to hundreds of captions on long content, VSeek settles at roughly 32 frames regardless of duration. VSeek-VETL samples slightly more than VSeek-EM because rewarding unverified claims less encourages more thorough retrieval, and that extra evidence is what closes the gap on long videos.



BibTeX

@inproceedings{goel2026incentivizing,
  author    = {Goel, Harsh and Sharan, SP and Shah, Sahil and Choi, Minkyu and An, Joungbin and Grauman, Kristen and Chinchali, Sandeep},
  title     = {Incentivizing Vision Language Models to Search for Long Video Question Answering},
  journal   = {Proceedings of the European Conference on Computer Vision (ECCV)},
  month     = {September},
  year      = {2026},
}