edition 2026-08 · published Aug 10, 2026
The nl2sql.ai Leaderboard — Launch Edition (August 2026)
| # | system | score | Δ |
|---|---|---|---|
| 1 | Cortex AnalystSnowflake | 88.0 | — |
| 2 | AI/BI GenieDatabricks | 86.0 | — |
| 3 | SQLCoder / DefogDefog.ai | 81.0 | — |
| 4 | VannaVanna.AI | 79.5 | — |
| 5 | Gemini in BigQueryGoogle Cloud | 78.0 | — |
| 6 | Wren AICanner | 76.5 | — |
| 7 | DB-GPTeosphoros-ai community | 74.0 | — |
| 8 | Copilot in Microsoft FabricMicrosoft | 72.5 | — |
methodology
Scores are 0–100, weighted across four pillars:
- Capability (40%) — quality of generated SQL on hard, messy schemas; benchmark evidence where public (BIRD, Spider 2.0); semantic grounding and self-correction.
- Trust & governance (25%) — how the system prevents wrong-but-plausible answers: semantic layers, verified queries, permissions awareness, auditability.
- Adoption signals (20%) — production usage, community momentum (stars, releases, case studies), ecosystem integrations.
- Openness (15%) — open weights/source, self-hosting, extensibility, transparent docs.
The launch edition was scored by the founding editorial pass and is the baseline. From here, the desk re-scores monthly: every score change must cite evidence (release notes, benchmark submissions, engineering posts), and every edition links its diff. Vendors never pay for placement; this publication has no commercial relationship with anything it ranks.