benchmark tracker
Reference points and reported results on the beat's two load-bearing evaluations. Every number links to its submission, paper, or leaderboard. Maintained continuously by the benchmark desk.
BIRD
Big Bench for Large-Scale Database Grounded Text-to-SQL Evaluation: 12,751 questions over 95 databases (33 GB) across 37 domains, emphasizing dirty values, external knowledge, and SQL efficiency. The de facto standard leaderboard for single-database text-to-SQL.
official leaderboard ↗Spider 2.0
Successor to Yale's Spider, built on real enterprise workflows: databases on Snowflake/BigQuery with 1,000+ column schemas, multiple dialects, and multi-step agentic tasks. Frontier models scored under 20% at launch — the current reality check for the field.
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| o1-preview agentic baseline (paper) | task success rate | 17.00 | Nov 1, 2024 | link ↗ |