The flavor of the bitter lesson for computer vision - Vincent Sitzmann
vincentsitzmann.com · 2,155 words · saved by 7 readers
Personal website
I believe that computer vision as we know it is about to go away. Historically, we have treated vision as a mapping from images to intermediate representations—classes, segmentation masks, or 3D reconstructions. But in the era of the Bitter Lesson, these distinct tasks are becoming qualitatively no different than edge detection: historical artifacts of scoping “solvable intermediate problems” rather than solving intelligence. While the “LLM moment” in NLP clarified that language modeling is the ultimate objective, the vision community is still debating the flavor of its own revolution. We cont
saved by
related reading
- World Action Model Atlasjoeclinton.me
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- General Instinct | Any frontier model. Any edge device.general-instinct.com
- The First Fully General Computer Action Model | blogsi.inc
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelarxiv.org
- Video models are zero-shot learners and reasonersarxiv.org
- A Functional Taxonomy of World Models - Dr. Fei-Fei Lidrfeifei.substack.com
- World Models: Computing the Uncomputablenotboring.co
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- World Models | Rohit Bandarurohitbandaru.github.io
- The Model That Dreams the Worldmoe-capital.com
- Explore | alphaXivalphaxiv.org