siboehm
siboehm.com · 1,050 words · saved by 3 readers
Simon Boehm's technical blog
siboehm This is the website of Simon Boehm. If you like these posts: Subscribe to the mailing list. Can Function Inlining Affect Floating Point Outputs? Exploring FMA and Other Consistency Issues At my job, I’m refactoring a 30k LOC codebase that simulates learning in the mammal brain. The emergent behavior of these large brain models is hard to test for, so we opted to take the safe route and preserve bit-equality in the weights of the trained model to guarantee that we are not breaking anything. The main downside of hash-based regression testing during refactoring is that it doesn’t check nu
saved by
related reading
- How to Parallelize a Transformer for Training — an explorable explanationezyang.github.io
- Big Boss (@0xBADB01E) on Xx.com
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- GitHub - wafer-ai/gpu-perf-engineering-resources: A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.github.com
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- KernelBench: Can LLMs Write GPU Kernels?scalingintelligence.stanford.edu
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- How To Scale Your Modeljax-ml.github.io
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com