flâneur

Introducing Supabase Evals

supabase.com · 1,108 words · saved by 1 readers

Our open-source benchmark for how well AI coding agents build with Supabase.

Today we're open sourcing supabase/evals, our benchmark and framework for testing how well AI agents build using Supabase. It runs coding agents including Claude Code, Codex, and OpenCode against real Supabase tasks, for example, building a schema, debugging a failed Edge Function, or fixing a broken RLS policy, and then scores how well they performed. It powers both our published benchmark and an internal regression suite we monitor daily. As more people ship Supabase projects through an agent instead of by hand, we wanted a way to measure that experience instead of guessing. The results…

saved by

related reading