✳flâneur — a map of the web's best reading
The Hardest Part of Shrinking a Robotics Model | Haptic Labs
hapticlabs.ai · 2,441 words · saved by 1 readers
Shrinking Pi0.5 by 2.9× on disk and 2.8× faster. The paper gives you the recipe; the pipeline tax is what makes it hard in production.
Your browser does not support the video tag. LIBERO-Spatial task 0, side by side: Pi0.5 FT teacher (4.14 B params, 8.3 GB) on the left, V8-trim + INT8 student (2.31 B params, 2.85 GB) on the right. Same task, ~3× smaller model. Shrinking Pi0.5 by 2.9× on disk and 2.8× faster. The paper gives you the recipe. We're sharing what it costs to actually run that recipe end-to-end. TL;DR Shrinking robotics models makes them faster and lets them run on cheaper, lower-power hardware. We shrank Pi0.5 using two techniques: layer pruning and INT8 quantization . End result: 3.0× smaller in memory, 2.9× smal
Explore this link on the map →related reading
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com
- Composer2.pdfcursor.com
- TurboQuant: Redefining AI efficiency with extreme compressionresearch.google
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- TurboQuant: Redefining AI efficiency with extreme compressionresearch.google
- A foolproof way to shrink deep learning models | MIT News | Massachusetts Institute of Technologynews.mit.edu
- [2309.10668] Language Modeling Is Compressionarxiv.org
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Studyarxiv.org
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Modelsarxiv.org