✳flâneur — a map of the web's best reading
siboehm
siboehm.com · 1,050 words · saved by 2 readers
Simon Boehm's technical blog
siboehm This is the website of Simon Boehm. If you like these posts: Subscribe to the mailing list. Can Function Inlining Affect Floating Point Outputs? Exploring FMA and Other Consistency Issues At my job, I’m refactoring a 30k LOC codebase that simulates learning in the mammal brain. The emergent behavior of these large brain models is hard to test for, so we opted to take the safe route and preserve bit-equality in the weights of the trained model to guarantee that we are not breaking anything. The main downside of hash-based regression testing during refactoring is that it doesn’t check nu
Explore this link on the map →saved by
related reading
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- How To Scale Your Modeljax-ml.github.io
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- Making Deep Learning go Brrrr From First Principleshorace.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- CVPR2023_eff_tutorial_molchanov.pdfnvlabs.github.io
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com