About Us - Colfax Research
Our team consists of mathematicians and scientists who bring formal academic training and deep analytical rigor to GPU kernel development. We have a demonstrated record of excellence across research, education, technical writing, and open-source contributions, pairing first principles reasoning about hardware, systems, and algorithms with hands-on kernel engineering. Research papers We publish research at the […]
Our team consists of mathematicians and scientists who bring formal academic training and deep analytical rigor to GPU kernel development. We have a demonstrated record of excellence across research, education, technical writing, and open-source contributions, pairing first principles reasoning about hardware, systems, and algorithms with hands-on kernel engineering. Research papers We publish research at the frontier of GPU performance and the mathematics that underpins it. Our publications include: FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision [blog,…
saved by
related reading
- KernelBench: Can LLMs Write GPU Kernels?scalingintelligence.stanford.edu
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklogsiboehm.com
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai
- We reverse-engineered Flash Attention 4modal.com
- GitHub - wafer-ai/gpu-perf-engineering-resources: A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.github.com
- Flash Attention from Scratch Part 1: Introlubits.ch
- Modern GPU Programming For MLSys — Modern GPU Programming For MLSysmlc.ai
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- How to Land a Frontier Lab Jobvladfeinberg.com
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernelsarxiv.org
- CUTLASS: Fast Linear Algebra in CUDA C++ | NVIDIA Technical Blogdeveloper.nvidia.com
- GPUs Go Brrr · Hazy Researchhazyresearch.stanford.edu