Kimi-K2.5 Inference Benchmark - Luminal
Luminal is an AI inference compiler that compiles and optimizes AI models for GPUs and ASICs, delivering the fastest, highest throughput inference in the world.
Kimi-K2.5 Inference Benchmark - Luminal Inference Benchmark moonshotai/Kimi-K2.5 on 8xH200 Model moonshotai/Kimi-K2.5 GPU 8xH200 Architecture 1T MoE, 32B active params Scenarios What Do the Scenarios Represent? Each benchmark scenario simulates a different real-world usage pattern with distinct input and output token profiles. Chatbot 1,024 in / 256 out A typical conversational turn: short context (prior messages plus a new prompt) and a moderate-length reply. This is the lightest workload and represents most chat-style applications. Prefill is fast, so TTFT is low and decode speed is at its p
Explore this link on the map →saved by
related reading
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Composer2.pdfcursor.com
- Optimizing inference · Hugging Facehuggingface.co
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Xiaomi MiMo, Explore and Lovemimo.xiaomi.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- A Guide to AI Inference Engineering - ByteByteGo Newsletterblog.bytebytego.com
- Inference.net | Full-Stack LLM Lifecycle Platforminference.net