How to Deploy Your Model
Model Gemma 4 31B Gemma 4 12B Gemma 4 26B A4B DeepSeek V4 Flash MXFP4/FP8 GLM 5.2 FP8 Kimi K2.6 INT4/BF16 Kimi K3 MXFP4/MXFP8 Muse Glimmer 30B Qwen3 4B BF16 Qwen3 4B FP8 Qwen3.8 27B Qwen3.6 35B A3B gpt-oss-120b MXFP4/BF16 gpt-oss-20b MXFP4/BF16 LLaMA 3 8B LLaMA 3 70B LLaMA 3.1 405B Workload Prefill tokens Decode tokens Decode batching ⓘ Maximum SLO-based Minimum (B=1) Minimum tok/s/user KV cache per sequence: 1.05 GB Hardware Nodes ⓘ 1 node 2 nodes 3 nodes 4 nodes 5 nodes 6 nodes 7 nodes 8 nodes Per-chip Sweep all sizes NVIDIA PRICE VS H100 Ⓘ RTX 4090 × H100 RTX 5090 × H100 A100 SXM × H100 RTX PRO 6000 × H100 H100 SXM × H100 H200 SXM × H100 B200 SXM × H100 GB200 NVL72 × H100 B300 SXM × H100 VR100 NVL72 × H100 AMD MI300X × H100 MI325X × H100 MI355X × H100 INTEL Gaudi 2 × H100 Gaudi 3 × H100 AWS Inferentia2 × H100 Trainium1 × H100 Trainium2 × H100 GOOGLE TPU v5e × H100 TPU v6e × H100 TPU v5p × H100 TPU v7x × H100 Overlap Memoryⓘ 90% Commsⓘ 65% Table Chart SLA filters 236/236 X Prefill to
42/86 #HardwareBest sharding by Requests/$Requests/$ ↓Prefill tok/chip/$Prefill tok/s/chipTTFTDecode tok/chip/$Decode tok/s/chipTok/s/userConfigs 1B200 SXM1 node✓DPA 2 × TP 41.17×1.03×23.5k22.9 ms1.21×5.52k359✓1⚠ 2A100 SXM1 node✓DPA 2 × TP 41.03×0.761×3.92k137 ms1.27×1.3k219✓1⚠ 3H100 SXM1 nodeHMVP✓DPA 2 × TP 41×0.918×10.5k51.2 ms0.991×2.25k379✓1⚠ 4H200 SXM1 node✓DPA 2 × TP 40.996×0.737×10.5k50.9 ms1.24×3.54k309✓1⚠ 5B300 SXM1 node✓DPA 80.78×0.758×26k26.5 ms0.737×5.03k252✓ 6GB200 NVL72————————0✓ 7RTX PRO 60001 node————————0✓ 8RTX 50901 nodeNo viable configurations———————0✓ 9RTX 40901…
saved by
related reading
- Together AI | The AI Native Cloudtogether.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- MatX: High-throughput chips for LLMsmatx.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- Replicate - Run AI with an APIreplicate.com
- Hugging Face – The AI community building the future.huggingface.co
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Open-Source Agentic Inference Benchmark | InferenceXinferencex.semianalysis.com
- Rent GPUs | Vast.aivast.ai
- Localmaxxing - Local LLM Inference Speed Testslocalmaxxing.com
- Compare AI Models: Pricing, Context & Benchmarks | OpenRouteropenrouter.ai