Spider 2.0 is the enterprise reality check
Real Snowflake and BigQuery projects, thousand-column schemas, multi-step tasks — and frontier models under 20% at launch.
If BIRD asked "can your model write correct SQL on messy data?", Spider 2.0 (announced November 2024) asks the question enterprises actually care about: can your agent do a data analyst's job?
The benchmark's tasks come from real enterprise-style workflows:
- Real platforms. Databases live on Snowflake, BigQuery, and friends — with their dialects, quirks, and permissions — not a local SQLite file.
- Real scale. Schemas run past 1,000 columns; context doesn't fit in a prompt, so systems must explore, search metadata, and read docs.
- Real workflows. Tasks chain steps — find the right tables, transform, aggregate, reconcile dialects — closer to a Jupyter session than a one-shot query.
The launch paper's agentic baseline built on o1-preview solved roughly 17% of tasks. Against Spider 1.0's saturated 91%+, that number reframed the entire field: the gap between benchmark success and deployable analyst agents was not a few points — it was most of the distance.
That is precisely what makes Spider 2.0 the leaderboard to watch now. It rewards the things production systems need — schema navigation at scale, dialect fluency, multi-step planning with self-correction — and it punishes demo-ware. Movement here is slower and noisier than BIRD's, which is the point: when a system posts a big Spider 2.0 jump, something architectural happened.
We track both boards on the benchmarks page, with every reported number linked to its submission or paper. When the first system crosses 50%, you'll read about it on the wire within the hour.
sources
- Spider 2.0 sitespider2-sql.github.io
- Spider 2.0 paper (arXiv 2411.07763)arxiv.org
- Welcome to nl2sql.aiAug 10, 2026
- The state of natural-language-to-SQL, August 2026Aug 10, 2026
- Why BIRD became the benchmark that mattersAug 10, 2026