Text-to-video retrieval has become essential as the volume of online video content surges, allowing users to locate relevant videos through natural language queries. Despite advancements in techniques such as dual encoders, attention models, and multimodal approaches over the last six years, significant questions persist regarding how these models perform, the impact of datasets, and the complexity of queries. In this research, we assess 14 leading retrieval techniques across three prominent datasets, applying a consistent preprocessing and evaluation strategy. We examine various aspects of captions, including their length, clarity, semantic categories, and the balance between action and scene descriptions, correlating these factors with model efficacy. Our findings indicate that concise and straightforward captions, particularly those that depict single actions or colors, yield better retrieval results. Conversely, more intricate events and detailed scene descriptions pose challenges for current models. Attention-based architectures excel with temporally complex queries, while dual-encoder and multimodal models are more effective with simpler captions. Furthermore, cross-dataset performance improves with larger and more varied caption collections, although generative captions do not always lead to better retrieval outcomes. This research underscores critical dataset considerations, benchmarking issues, and the relationship between query types and model structures, offering insights for enhancing text-to-video retrieval systems.
Exploring the Challenges of Text-to-Video Retrieval Performance
This study investigates the intricacies of text-to-video retrieval, focusing on model performance and dataset characteristics.
