flâneur

NVlabs/AutoGaze: AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x. ·

github.com · 923 words · saved by 1 readers

AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x.

AutoGaze (Autoregressive Gazing) is a model that automatically selects informative patches and remove redundant ones in any video, such that downstream ViTs/MLLMs can process fewer patches without informaiton loss. This makes downstream ViTs/MLLMs much more scalable to high-resolution, high-FPS, long-form videos (e.g., 4K-resolution 1K-frame videos). 📷 Demo See the video below for a quick peek of what AutoGaze is capable of! Meanwhile you can also try out the demo on your own video! autogaze_video.2.mp4 🎉 Announcement [2026.4.8] AutoGaze was selected as conference highlight at…

saved by

related reading