live wire
nl2sql.ai is live — first leaderboard edition published, directory and benchmark tracker onlinenl2sql.aiBIRD leaderboard: top cluster holds in the low-to-mid 70s on dev execution accuracy; human reference 92.96bird-bench.github.ioSpider 2.0 remains the wall: launch-paper agentic baseline solved ~17% of enterprise tasksarXivDirectory day one: 12 systems catalogued across cloud-native, OSS, and research categoriesnl2sql.aiVanna remains the most-starred OSS text-to-SQL framework; RAG-on-your-own-pairs still the default patternGitHubWren AI ships steadily on its MDL semantic layer — the OSS counterpart to vendor semantic modelsGitHubUber QueryGPT post remains the canonical enterprise-scale case study: routing beats generation at 1000s of tablesUber EngineeringWatch item: semantic-layer interop — every platform has one, none of them talk to each otheranalysis deskDB-GPT community keeps shipping: fine-tuning hub and AWEL workflows anchor the self-hosted stackGitHubDesk assignments filed: releases and papers, leaderboards, the directory. Cadence: continuousnl2sql.ainl2sql.ai is live — first leaderboard edition published, directory and benchmark tracker onlinenl2sql.aiBIRD leaderboard: top cluster holds in the low-to-mid 70s on dev execution accuracy; human reference 92.96bird-bench.github.ioSpider 2.0 remains the wall: launch-paper agentic baseline solved ~17% of enterprise tasksarXivDirectory day one: 12 systems catalogued across cloud-native, OSS, and research categoriesnl2sql.aiVanna remains the most-starred OSS text-to-SQL framework; RAG-on-your-own-pairs still the default patternGitHubWren AI ships steadily on its MDL semantic layer — the OSS counterpart to vendor semantic modelsGitHubUber QueryGPT post remains the canonical enterprise-scale case study: routing beats generation at 1000s of tablesUber EngineeringWatch item: semantic-layer interop — every platform has one, none of them talk to each otheranalysis deskDB-GPT community keeps shipping: fine-tuning hub and AWEL workflows anchor the self-hosted stackGitHubDesk assignments filed: releases and papers, leaderboards, the directory. Cadence: continuousnl2sql.ai
nl2sql.ai

edition 2026-08 · published Aug 10, 2026

The nl2sql.ai Leaderboard — Launch Edition (August 2026)

#systemscoreΔ
1Cortex AnalystSnowflake88.0
2AI/BI GenieDatabricks86.0
3SQLCoder / DefogDefog.ai81.0
4VannaVanna.AI79.5
5Gemini in BigQueryGoogle Cloud78.0
6Wren AICanner76.5
7DB-GPTeosphoros-ai community74.0
8Copilot in Microsoft FabricMicrosoft72.5
methodology

Scores are 0–100, weighted across four pillars:

  • Capability (40%) — quality of generated SQL on hard, messy schemas; benchmark evidence where public (BIRD, Spider 2.0); semantic grounding and self-correction.
  • Trust & governance (25%) — how the system prevents wrong-but-plausible answers: semantic layers, verified queries, permissions awareness, auditability.
  • Adoption signals (20%) — production usage, community momentum (stars, releases, case studies), ecosystem integrations.
  • Openness (15%) — open weights/source, self-hosting, extensibility, transparent docs.

The launch edition was scored by the founding editorial pass and is the baseline. From here, the desk re-scores monthly: every score change must cite evidence (release notes, benchmark submissions, engineering posts), and every edition links its diff. Vendors never pay for placement; this publication has no commercial relationship with anything it ranks.