Localmaxxing - Local LLM Inference Benchmarks
Community-sourced benchmarks for local LLM inference. Compare tokens/sec across GPUs, Apple Silicon, CPUs, and engines like llama.cpp, Ollama, and LM Studio.
LocalMaxxing Speed-test your local LLM rig.Compare it with the world. Community speed tests for local LLM inference. Track speed, compare hardware, and find your optimal setup. Every number on this site comes from a community-submitted run on real hardware — no vendor speed tests. Explore the platform Live from the community Benchmark scoreboard Browse all benchmarks → Fresh on the marketplace Browse marketplace → How it works 01 Sign in & create a key Sign in with GitHub, then create an API key in your dashboard so the CLI and your agents can submit runs. 02 Run a speed test…
saved by
related reading
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- LLM Visualizationbbycroft.net
- Prompting best practicesdocs.anthropic.com
- Together AI | The AI Native Cloudtogether.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- LocalScore - Local AI Benchmarklocalscore.ai
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- Open-Source Agentic Inference Benchmark | InferenceXinferencex.semianalysis.com
- Perplexityperplexity.ai
- BalatroBenchbalatrobench.com
- Datacurve | The data engine for frontier AIdatacurve.ai
- Compare AI Models: Pricing, Context & Benchmarks | OpenRouteropenrouter.ai