Auto-research with codex: How I achieved a 232x Faster Kernel over baseline with Codex in GPU Mode's qr_v2 problem – sankalp's blog
Table of Contents Intro Contest in short Problem intro Why this problem is auto-research-able Learning Enough to Ask Better Questions (Optional) Math f...
08 Jul, 2026 Table of Contents Intro Contest in short Problem intro Why this problem is auto-research-able Learning Enough to Ask Better Questions (Optional) Math for QR decomposition: Householder reflections Make serial work small with the help of the blocked Householder algorithm Other challenges Codex-maxxing Kernel progress breakthroughs Breakthrough ideas Introducing idea diversity to escape the local maxima Implementation Hints What I could have done better Conclusion References Acknowledgements Intro Contest in short GPU Mode, in collab with Core Automation,…
saved by
related reading
- QR Decomp at the Speed of Lightml-mike.com
- GitHub - wafer-ai/gpu-perf-engineering-resources: A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.github.com
- Decoding Speculative Decoding from First Principlesjwlabs.vercel.app
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- KernelBench: Can LLMs Write GPU Kernels?scalingintelligence.stanford.edu
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- siboehmsiboehm.com
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Kevin-32B: Multi-Turn RL for Writing CUDA Kernels | Cognitioncognition.ai
- About Us - Colfax Researchresearch.colfax-intl.com
- KernelBench v0.1 | Scaling Intelligence Lab at Stanford Universityscalingintelligence.stanford.edu