flâneur

AI Inference Benchmarks | InferenceX by SemiAnalysis

inferencex.semianalysis.com · 193 words · saved by 1 readers

Compare AI inference latency, throughput, and time-to-first-token across GPUs and providers. Real benchmarks on NVIDIA GB200, H100, AMD MI355X, and more.

AgentX / live results Compare Realistic Agentic Inference Perf Long Context Multi Turn Inference Performance. Compare Across Google TPUv7 Ironwood, OpenAI Jalapeño, MI355X, GB300 NVL72, GB200 NVL72, B200, H200, H100, RTX Pro, and soon TPUv8i, SambaNova SN50, Rubin NVL72, MI455X UALoE72 Token Revenue CalculatorDashboard Every Result Is Transparently done through Public GitHub Actions Automation Every data point on the dashboard is produced by a public GitHub Actions workflow run. The recipe lives in the repo, the run executes on the actual target hardware, and the full logs and artifacts…

saved by

related reading