[2404.01476] TraveLER: A Multi-LMM Agent Framework for Video Question-Answering
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
View PDF HTML (experimental) Abstract:Recently, image-based Large Multimodal Models (LMMs) have made significant progress in video question-answering (VideoQA) using a frame-wise approach by leveraging large-scale pretraining in a zero-shot manner. Nevertheless, these models need to be capable of finding relevant information, extracting it, and answering the question simultaneously. Currently, existing methods perform all of these steps in a single pass without being able to adapt if insufficient or incorrect information is collected. To overcome this, we introduce a modular multi-LMM agent…
saved by
related reading
- Explore | alphaXivalphaxiv.org
- 2501.05874arxiv.org
- Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazingarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- 2403.09611.pdfarxiv.org
- Video models are zero-shot learners and reasonersarxiv.org
- Less Detail, Better Answers: Degradation-Driven Prompting for VQAarxiv.org
- Kimi K2.5: Visual Agentic Intelligence1kpapers.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- [2508.16201] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruningarxiv.org
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelarxiv.org
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org