flâneur — a map of the web's best reading

Yann LeCun on X: "A short post on the best architectures for real-time image and video processing. TL;DR: use convolutions with stride or pooling at the low levels, and stick self-attention circuits at higher levels, where feature vectors represent objects. PS: ready to bet that Tesla FSD uses" / X

x.com · saved by 1 readers

To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore Notifications Messages Grok Lists Bookmarks Communities Premium Profile More Post Djsshsh @Djsshsh8 Post See new posts Conversation Yann LeCun @ylecun A short post on the best architectures for real-time image and video processing. TL;DR: use convolutions with stride or pooling at the low levels, and stick self-attention circuits at higher levels, where feature vectors represent objects. PS: ready to bet that Tesla FSD uses convolutions (or perhaps more complex *local* operators) at the low levels, combined with more global circuits at higher levels (perhaps using self-attention). Transformers on low-level patch embeddings are a complete waste of electrons. Quote Yann LeCun @ylecun · May 30 Replying to @___Harald___ I'm not saying ViTs are not practical (we use them). I'm saying they are way too slow and inefficient to be practical for real-time processing of high-resolution images and video. [Also,

Explore this link on the map →

saved by