The challenge of retrieving videos from large datasets often stems from vague user queries, which traditional single-round retrieval methods struggle to address. These methods typically encounter limitations due to their inability to incorporate effective feedback mechanisms, leading to what is termed the 'Intent-Query Gap.' This gap arises when a user's true intent is not adequately represented by a straightforward text query. To tackle this issue, we present the ADEPT framework, an innovative agent that operates without the need for prior training. ADEPT utilizes an entropy-driven decision-making process that adeptly navigates dialogue by alternating between two strategies: ASK and REFINE. Through rigorous testing on two demanding datasets, ADEPT has shown remarkable improvements over existing non-interactive, heuristic, and Video-LLM approaches. The primary achievement of this research lies in establishing a clear and efficient interactive strategy that enhances interpretability and sets a new standard for performance in the domain of interactive video retrieval.
ADEPT: A Novel Approach to Interactive Video Retrieval
ADEPT introduces a groundbreaking method for improving video retrieval from extensive datasets by addressing user query ambiguities.
