Startups · Aug 3
Supabase Evals Is a Bet That Every Developer Platform Needs Its Own Agent Benchmark
The open source framework tests Claude Code, Codex, and OpenCode on real Supabase work, then feeds a public benchmark and daily regression suite.
- Treat agent evals as product QA, not model theater.
- Benchmark the workflows users actually run, then wire results into daily regression checks.
- Platform specific tests can expose failures generic coding benchmarks miss.