Xianbang Wang — research & writing
peppaking8.github.io · 2,187 words · saved by 2 readers
Xianbang Wang — MIT '29, deep learning · vision · RL. Personal site and technical blog.
A practical story about writing TPU Pallas kernels: from naive attention to native backward kernels, sparse attention, and a 1.57x nanoGPT training speedup. GPU users have CUDA and Triton tutorials everywhere; TPU users usually hear the simpler advice: write JAX and trust XLA. This post asks when that advice is still leaving performance on the table. We build a causal attention Pallas kernel, add the native backward pass needed for training, fuse the Q/K preprocessing and layout boundary, and integrate the result into a nanoGPT-style language model. The final v6e-16 training run improves from
saved by
related reading
- LLM Visualizationbbycroft.net
- Tristan's Site - Tristan Humethume.ca
- Mingxuan Yanwaterhyacinthinnanhu.github.io
- Amelia Wattenbergerwattenberger.com
- Black Forest Labs - Frontier AI Labbfl.ai
- Replicate - Run AI with an APIreplicate.com
- Yufeng Zhaoyufengzhao.com
- Explore | alphaXivalphaxiv.org
- justinzwu.comjustinzwu.com
- Cosmoscosmos.so
- GitHub - labmlai/annotated_deep_learning_paper_implementations: 🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adagithub.com
- Lummi — Free AI Stock Images, Illustrations & 3Dlummi.ai