✳flâneur — a map of the web's best reading
The flavor of the bitter lesson for computer vision - Vincent Sitzmann
vincentsitzmann.com · 2,155 words · saved by 7 readers
Personal website
I believe that computer vision as we know it is about to go away. Historically, we have treated vision as a mapping from images to intermediate representations—classes, segmentation masks, or 3D reconstructions. But in the era of the Bitter Lesson, these distinct tasks are becoming qualitatively no different than edge detection: historical artifacts of scoping “solvable intermediate problems” rather than solving intelligence. While the “LLM moment” in NLP clarified that language modeling is the ultimate objective, the vision community is still debating the flavor of its own revolution. We cont
Explore this link on the map →saved by
related reading
- The First Fully General Computer Action Model | blogsi.inc
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- Video models are zero-shot learners and reasonersarxiv.org
- World Models: Computing the Uncomputablenotboring.co
- A Functional Taxonomy of World Models - Dr. Fei-Fei Lidrfeifei.substack.com
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- The Model That Dreams the Worldmoe-capital.com
- [2602.06001] Visuo-Tactile World Modelsarxiv.org
- World Models | Rohit Bandarurohitbandaru.github.io
- State of Robot Learning, December 2025vedder.io
- 3D as code | World Labsworldlabs.ai
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai