✳flâneur — a map of the web's best reading
Latency vs. Throughput in Machine Learning Pipelines
simonwenkel.com · 653 words · saved by 1 readers
Providing practical tutorials and unconventional views on AI for physical world applications.
Contents Introduction Problem-centric View Implementation Complexity View Systems Configuration Introduction “Latency vs. Throughput in machine learning pipelines” is topic I end up discussing almost as often as how to avoid common mistakes . While many people seem to primarily think about the training side of pipelines, I’d consider the inference side a lot more interesting and challenging as there are many more trade-offs. IMHO, throughput is key on the training side - latency basically doesn’t matter although unnecessarily high latency decreases throughput of course. The thoughts presented
Explore this link on the map →related reading
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Making Deep Learning go Brrrr From First Principleshorace.io
- How To Scale Your Modeljax-ml.github.io
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- LLM Engineer's Almanac - Workloads | Modalmodal.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Kimi-K2.5 Inference Benchmark - Luminalluminal.com
- How is LLaMa.cpp possible?finbarr.ca
- Unlocking the full power of NVIDIA H100 GPUs for ML inference with TensorRTbaseten.co
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu