NVlabs/AutoGaze: AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x. ·
github.com · 923 words · saved by 1 readers
AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x.
AutoGaze (Autoregressive Gazing) is a model that automatically selects informative patches and remove redundant ones in any video, such that downstream ViTs/MLLMs can process fewer patches without informaiton loss. This makes downstream ViTs/MLLMs much more scalable to high-resolution, high-FPS, long-form videos (e.g., 4K-resolution 1K-frame videos). 📷 Demo See the video below for a quick peek of what AutoGaze is capable of! Meanwhile you can also try out the demo on your own video! autogaze_video.2.mp4 🎉 Announcement [2026.4.8] AutoGaze was selected as conference highlight at…
saved by
related reading
- Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazingarxiv.org
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- Trending Papers - Hugging Facepaperswithcode.com
- Together AI | The AI Native Cloudtogether.ai
- Replicate - Run AI with an APIreplicate.com
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automaticallygithub.com
- Video models are zero-shot learners and reasonersarxiv.org
- francesco215.github.io/autoregressive_diffusion/francesco215.github.io
- TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answeringarxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- [2404.02905] Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Predictionarxiv.org
- A Dive into Text-to-Video Modelshuggingface.co