✳flâneur — a map of the web's best reading
LIME YAO
0 followers · 403 views
Open this reading profile →
on the atlas — 22
- A Parallel Computing Primer1 savers
- The 2026 AI Index Report | Stanford HAI1 savers
- Worklog: Optimising GEMM on NVIDIA H100 for cuBLAS-like Performance (WIP) | Hamza's Blog1 savers
- Katakana Writing Practice | Characters | Japanese-Lesson.com1 savers
- Hiragana Writing Practice | Characters | Japanese-Lesson.com1 savers
- A Deep Dive into LLM Inference Latencies1 savers
- KV cache strategies1 savers
- API Reference — TensorRT LLM1 savers
- Efficient GEMM in CUDA — NVIDIA CUTLASS Documentation1 savers
- NVIDIA Tensor Core Evolution: From Volta To Blackwell1 savers
- CVPR2023_eff_tutorial_molchanov.pdf1 savers
- Formula for Human Genius and Creativity1 savers
- Another Giant Leap: The Rubin CPX Specialized Accelerator & Rack – SemiAnalysis1 savers
- BIST for Analog Weenies | Analog Devices1 savers
- Python API reference — nvMatmulHeuristics1 savers
- H100 vs GB200 NVL72 Training Benchmarks – Power, TCO, and Reliability Analysis, Software Improvement Over Time – SemiAnalysis1 savers
- How to Optimize a CUDA Matmul Kernel for cuBLAS-like Performance: a Worklog23 savers
- The Smol Training Playbook: The Secrets to Building World-Class LLMs - a Hugging Face Space by HuggingFaceTB13 savers
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordić9 savers
- All About Rooflines | How To Scale Your Model9 savers
- How to Think About TPUs | How To Scale Your Model5 savers
- GPU Performance Background User's Guide - NVIDIA Docs2 savers