Inferact
Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress.
Today, we're proud to announce Inferact, a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. The Challenge Inference is not solved. It's getting harder. Models grow larger. New architectures proliferate: mixture-of-experts, multimodal, agentic. Every breakthrough demands new infrastructure. Meanwhile, hardware fragments: more accelerators, more programming models, and more combinations to optimize.…
saved by
related reading
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- Together AI | The AI Native Cloudtogether.ai
- Inference.net | Full-Stack LLM Lifecycle Platforminference.net
- As Rocks May Think | Eric Jangevjang.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- LLM Inference Handbookhandbook.modular.com
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Optimizing inference · Hugging Facehuggingface.co
- 2025: The year in LLMssimonwillison.net
- Things we learned about LLMs in 2024simonwillison.net
- vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention | vLLM Blogblog.vllm.ai
- Introduction to vLLM and PagedAttentionblog.runpod.io