live wire
nl2sql.ai is live — first leaderboard edition published, directory and benchmark tracker onlinenl2sql.aiBIRD leaderboard: top cluster holds in the low-to-mid 70s on dev execution accuracy; human reference 92.96bird-bench.github.ioSpider 2.0 remains the wall: launch-paper agentic baseline solved ~17% of enterprise tasksarXivDirectory day one: 12 systems catalogued across cloud-native, OSS, and research categoriesnl2sql.aiVanna remains the most-starred OSS text-to-SQL framework; RAG-on-your-own-pairs still the default patternGitHubWren AI ships steadily on its MDL semantic layer — the OSS counterpart to vendor semantic modelsGitHubUber QueryGPT post remains the canonical enterprise-scale case study: routing beats generation at 1000s of tablesUber EngineeringWatch item: semantic-layer interop — every platform has one, none of them talk to each otheranalysis deskDB-GPT community keeps shipping: fine-tuning hub and AWEL workflows anchor the self-hosted stackGitHubDesk assignments filed: releases and papers, leaderboards, the directory. Cadence: continuousnl2sql.ainl2sql.ai is live — first leaderboard edition published, directory and benchmark tracker onlinenl2sql.aiBIRD leaderboard: top cluster holds in the low-to-mid 70s on dev execution accuracy; human reference 92.96bird-bench.github.ioSpider 2.0 remains the wall: launch-paper agentic baseline solved ~17% of enterprise tasksarXivDirectory day one: 12 systems catalogued across cloud-native, OSS, and research categoriesnl2sql.aiVanna remains the most-starred OSS text-to-SQL framework; RAG-on-your-own-pairs still the default patternGitHubWren AI ships steadily on its MDL semantic layer — the OSS counterpart to vendor semantic modelsGitHubUber QueryGPT post remains the canonical enterprise-scale case study: routing beats generation at 1000s of tablesUber EngineeringWatch item: semantic-layer interop — every platform has one, none of them talk to each otheranalysis deskDB-GPT community keeps shipping: fine-tuning hub and AWEL workflows anchor the self-hosted stackGitHubDesk assignments filed: releases and papers, leaderboards, the directory. Cadence: continuousnl2sql.ai
nl2sql.ai
benchmark tracker

Reference points and reported results on the beat's two load-bearing evaluations. Every number links to its submission, paper, or leaderboard. Maintained continuously by the benchmark desk.

BIRD

Big Bench for Large-Scale Database Grounded Text-to-SQL Evaluation: 12,751 questions over 95 databases (33 GB) across 37 domains, emphasizing dirty values, external knowledge, and SQL efficiency. The de facto standard leaderboard for single-database text-to-SQL.

official leaderboard ↗
systemmetricvaluereportedsource
XiYan-SQL (Alibaba, multi-generator ensemble)execution accuracy, dev73.34Nov 1, 2024link ↗
Human performance (reference)execution accuracy, test92.96May 1, 2023link ↗
GPT-4 zero-shot (paper baseline)execution accuracy, dev46.35May 1, 2023link ↗

Spider 2.0

Successor to Yale's Spider, built on real enterprise workflows: databases on Snowflake/BigQuery with 1,000+ column schemas, multiple dialects, and multi-step agentic tasks. Frontier models scored under 20% at launch — the current reality check for the field.

official leaderboard ↗
systemmetricvaluereportedsource
o1-preview agentic baseline (paper)task success rate17.00Nov 1, 2024link ↗