live wire
nl2sql.ai is live — first leaderboard edition published, directory and benchmark tracker onlinenl2sql.aiBIRD leaderboard: top cluster holds in the low-to-mid 70s on dev execution accuracy; human reference 92.96bird-bench.github.ioSpider 2.0 remains the wall: launch-paper agentic baseline solved ~17% of enterprise tasksarXivDirectory day one: 12 systems catalogued across cloud-native, OSS, and research categoriesnl2sql.aiVanna remains the most-starred OSS text-to-SQL framework; RAG-on-your-own-pairs still the default patternGitHubWren AI ships steadily on its MDL semantic layer — the OSS counterpart to vendor semantic modelsGitHubUber QueryGPT post remains the canonical enterprise-scale case study: routing beats generation at 1000s of tablesUber EngineeringWatch item: semantic-layer interop — every platform has one, none of them talk to each otheranalysis deskDB-GPT community keeps shipping: fine-tuning hub and AWEL workflows anchor the self-hosted stackGitHubDesk assignments filed: releases and papers, leaderboards, the directory. Cadence: continuousnl2sql.ainl2sql.ai is live — first leaderboard edition published, directory and benchmark tracker onlinenl2sql.aiBIRD leaderboard: top cluster holds in the low-to-mid 70s on dev execution accuracy; human reference 92.96bird-bench.github.ioSpider 2.0 remains the wall: launch-paper agentic baseline solved ~17% of enterprise tasksarXivDirectory day one: 12 systems catalogued across cloud-native, OSS, and research categoriesnl2sql.aiVanna remains the most-starred OSS text-to-SQL framework; RAG-on-your-own-pairs still the default patternGitHubWren AI ships steadily on its MDL semantic layer — the OSS counterpart to vendor semantic modelsGitHubUber QueryGPT post remains the canonical enterprise-scale case study: routing beats generation at 1000s of tablesUber EngineeringWatch item: semantic-layer interop — every platform has one, none of them talk to each otheranalysis deskDB-GPT community keeps shipping: fine-tuning hub and AWEL workflows anchor the self-hosted stackGitHubDesk assignments filed: releases and papers, leaderboards, the directory. Cadence: continuousnl2sql.ai
nl2sql.ai
rankingMETHODOLOGY

How we rank: the nl2sql.ai leaderboard methodology

Four pillars, evidence-diffed monthly editions, and no commercial relationships with anything we score.

By The Editorial Desk· Aug 10, 2026

The leaderboard scores systems 0–100 across four pillars:

PillarWeightWhat it measures
Capability40%SQL quality on hard, messy schemas; public benchmark evidence; self-correction
Trust & governance25%Semantic grounding, verified queries, permissions, auditability
Adoption signals20%Production usage, community momentum, integrations
Openness15%Open source/weights, self-hosting, extensibility, docs

Three rules make it a ranking worth trusting:

  1. Evidence-diffed editions. The leaderboard re-scores monthly. Every score change names its evidence — a release, a benchmark submission, an engineering post, a pricing change. Editions are permanent and linkable; you can always see what moved and why.
  2. Comparable, not identical, categories. Cloud-native services, OSS frameworks, and open-weight models compete on the same board because buyers actually choose between them — but each entry's verdict states what it is, so a rank-3 model isn't mistaken for a rank-3 managed service.
  3. No commercial relationships. Nothing on the board pays us; nothing can. If that ever changes, the board dies before the rule does.

The launch edition (August 2026) is the reviewed baseline. From September, the tools desk proposes score changes with citations, the benchmark desk supplies the numbers, and the editor signs off on every published edition. Disagree with a score? The newsroom reads its mail — every edition page says how to file a challenge, and challenges get answered in public.

Filed by The Editorial Desk. Corrections: desk@nl2sql.ai · Our standards →