BalatroBench
balatrobench.com · 124 words · saved by 1 readers
Leaderboard benchmarking LLMs playing Balatro: rounds, tool-call reliability, cost, and speed.
Click on a row to explore individual model runs Visit on a desktop for the full interactive experience Model leaderboard with rounds, tool-call reliability, tokens, time, and cost # Model Vendor Round Average final round reached across all runs (± std. dev.). Responses with valid tool calls that can be executed in the current game state. Responses with valid tool calls that cannot be executed in the current game state. Responses without valid tool calls. In / Average input tokens per tool call (± std. dev.). Out / Average output tokens per tool call (± std. dev., including…
saved by
related reading
- BalatroBenchbalatrobench.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- LLM Visualizationbbycroft.net
- Perplexityperplexity.ai
- Datacurve | The data engine for frontier AIdatacurve.ai
- There's An AI For That® — The front page of AItheresanaiforthat.com
- Localmaxxing - Local LLM Inference Speed Testslocalmaxxing.com
- Lakera – Test your AI hacking skillsgandalf.lakera.ai
- Compare AI Models: Pricing, Context & Benchmarks | OpenRouteropenrouter.ai
- OpenAI | Research & Deploymentopenai.com
- Prompting best practicesdocs.anthropic.com