rankingMETHODOLOGY
How we rank: the nl2sql.ai leaderboard methodology
Four pillars, evidence-diffed monthly editions, and no commercial relationships with anything we score.
By The Editorial Desk· Aug 10, 2026
The leaderboard scores systems 0–100 across four pillars:
| Pillar | Weight | What it measures |
|---|---|---|
| Capability | 40% | SQL quality on hard, messy schemas; public benchmark evidence; self-correction |
| Trust & governance | 25% | Semantic grounding, verified queries, permissions, auditability |
| Adoption signals | 20% | Production usage, community momentum, integrations |
| Openness | 15% | Open source/weights, self-hosting, extensibility, docs |
Three rules make it a ranking worth trusting:
- Evidence-diffed editions. The leaderboard re-scores monthly. Every score change names its evidence — a release, a benchmark submission, an engineering post, a pricing change. Editions are permanent and linkable; you can always see what moved and why.
- Comparable, not identical, categories. Cloud-native services, OSS frameworks, and open-weight models compete on the same board because buyers actually choose between them — but each entry's verdict states what it is, so a rank-3 model isn't mistaken for a rank-3 managed service.
- No commercial relationships. Nothing on the board pays us; nothing can. If that ever changes, the board dies before the rule does.
The launch edition (August 2026) is the reviewed baseline. From September, the tools desk proposes score changes with citations, the benchmark desk supplies the numbers, and the editor signs off on every published edition. Disagree with a score? The newsroom reads its mail — every edition page says how to file a challenge, and challenges get answered in public.
Filed by The Editorial Desk. Corrections: desk@nl2sql.ai · Our standards →
more from the desk
- Welcome to nl2sql.aiAug 10, 2026
- The state of natural-language-to-SQL, August 2026Aug 10, 2026
- Why BIRD became the benchmark that mattersAug 10, 2026