live wire
nl2sql.ai is live — first leaderboard edition published, directory and benchmark tracker onlinenl2sql.aiBIRD leaderboard: top cluster holds in the low-to-mid 70s on dev execution accuracy; human reference 92.96bird-bench.github.ioSpider 2.0 remains the wall: launch-paper agentic baseline solved ~17% of enterprise tasksarXivDirectory day one: 12 systems catalogued across cloud-native, OSS, and research categoriesnl2sql.aiVanna remains the most-starred OSS text-to-SQL framework; RAG-on-your-own-pairs still the default patternGitHubWren AI ships steadily on its MDL semantic layer — the OSS counterpart to vendor semantic modelsGitHubUber QueryGPT post remains the canonical enterprise-scale case study: routing beats generation at 1000s of tablesUber EngineeringWatch item: semantic-layer interop — every platform has one, none of them talk to each otheranalysis deskDB-GPT community keeps shipping: fine-tuning hub and AWEL workflows anchor the self-hosted stackGitHubDesk assignments filed: releases and papers, leaderboards, the directory. Cadence: continuousnl2sql.ainl2sql.ai is live — first leaderboard edition published, directory and benchmark tracker onlinenl2sql.aiBIRD leaderboard: top cluster holds in the low-to-mid 70s on dev execution accuracy; human reference 92.96bird-bench.github.ioSpider 2.0 remains the wall: launch-paper agentic baseline solved ~17% of enterprise tasksarXivDirectory day one: 12 systems catalogued across cloud-native, OSS, and research categoriesnl2sql.aiVanna remains the most-starred OSS text-to-SQL framework; RAG-on-your-own-pairs still the default patternGitHubWren AI ships steadily on its MDL semantic layer — the OSS counterpart to vendor semantic modelsGitHubUber QueryGPT post remains the canonical enterprise-scale case study: routing beats generation at 1000s of tablesUber EngineeringWatch item: semantic-layer interop — every platform has one, none of them talk to each otheranalysis deskDB-GPT community keeps shipping: fine-tuning hub and AWEL workflows anchor the self-hosted stackGitHubDesk assignments filed: releases and papers, leaderboards, the directory. Cadence: continuousnl2sql.ai
nl2sql.ai
benchmarkBENCHMARK DESK

Spider 2.0 is the enterprise reality check

Real Snowflake and BigQuery projects, thousand-column schemas, multi-step tasks — and frontier models under 20% at launch.

By The Benchmark Desk· Aug 10, 2026

If BIRD asked "can your model write correct SQL on messy data?", Spider 2.0 (announced November 2024) asks the question enterprises actually care about: can your agent do a data analyst's job?

The benchmark's tasks come from real enterprise-style workflows:

  • Real platforms. Databases live on Snowflake, BigQuery, and friends — with their dialects, quirks, and permissions — not a local SQLite file.
  • Real scale. Schemas run past 1,000 columns; context doesn't fit in a prompt, so systems must explore, search metadata, and read docs.
  • Real workflows. Tasks chain steps — find the right tables, transform, aggregate, reconcile dialects — closer to a Jupyter session than a one-shot query.

The launch paper's agentic baseline built on o1-preview solved roughly 17% of tasks. Against Spider 1.0's saturated 91%+, that number reframed the entire field: the gap between benchmark success and deployable analyst agents was not a few points — it was most of the distance.

That is precisely what makes Spider 2.0 the leaderboard to watch now. It rewards the things production systems need — schema navigation at scale, dialect fluency, multi-step planning with self-correction — and it punishes demo-ware. Movement here is slower and noisier than BIRD's, which is the point: when a system posts a big Spider 2.0 jump, something architectural happened.

We track both boards on the benchmarks page, with every reported number linked to its submission or paper. When the first system crosses 50%, you'll read about it on the wire within the hour.

sources

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →