✳flâneur — a map of the web's best reading
Lei.Chat()
lei.chat · 435 words · saved by 1 readers
© 2018 - 2026 Lei Zhang · Powered by the Eureka theme for Hugo
Lei.Chat() Lei Zhang Director of Engineering AMD AI Group AI Compiler & Runtime. Currently: Triton, IREE, MLIR, LLVM. Previously: SPIR-V, Vulkan, Metal. Recent Posts Gluon: Explicit Performance Gluon enhances the Triton language and compiler solutions with an additional approach towards GPU kernel programming. It strikes a different balance in the portability and performance spectrum to expose more compiler internals; thus giving developers more explicit controls to reach higher performance ceiling. In this blog post I’ll explain Gluon per my understanding. I will also use this as an opportuni
Explore this link on the map →saved by
related reading
- GitHub - triton-lang/triton: Development repository for the Triton language and compiler · GitHubgithub.com
- A friendly introduction to machine learning compilers and optimizershuyenchip.com
- How To Scale Your Modeljax-ml.github.io
- TPU Deep Divehenryhmko.github.io
- [2410.20399] ThunderKittens: Simple, Fast, and Adorable AI Kernelsarxiv.org
- GitHub - linkedin/Liger-Kernel: Efficient Triton Kernels for LLM Training · GitHubgithub.com
- Tristan's Site - Tristan Humethume.ca
- Triton: An Intermediate Language and Compiler for Tiled Neural Network Computationseecs.harvard.edu
- Composer2.pdfcursor.com
- Add GPU support to ggml · ggml-org/llama.cpp · Discussion #915 · GitHubgithub.com
- State of torch.compile for training (August 2025) : ezyang's blogblog.ezyang.com
- PyTorch internals : ezyang's blogblog.ezyang.com