Kimi-K2.5 Inference Benchmark - Luminal
Luminal is an AI inference compiler that compiles and optimizes AI models for GPUs and ASICs, delivering the fastest, highest throughput inference in the world.
Kimi-K2.5 Inference Benchmark - Luminal Inference Benchmark moonshotai/Kimi-K2.5 on 8xH200 Model moonshotai/Kimi-K2.5 GPU 8xH200 Architecture 1T MoE, 32B active params Scenarios What Do the Scenarios Represent? Each benchmark scenario simulates a different real-world usage pattern with distinct input and output token profiles. Chatbot 1,024 in / 256 out A typical conversational turn: short context (prior messages plus a new prompt) and a moderate-length reply. This is the lightest workload and represents most chat-style applications. Prefill is fast, so TTFT is low and decode speed is at its p
saved by
related reading
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)blog.kog.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Together AI | The AI Native Cloudtogether.ai
- Open-Source Agentic Inference Benchmark | InferenceXinferencex.semianalysis.com
- Composer2.pdfcursor.com
- Optimizing inference · Hugging Facehuggingface.co
- LLM Engineer's Almanac - Advisormodal.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- LLM Inference Handbookhandbook.modular.com