Putting Task Expertise into RL Achieves State-of-the-Art Performance on Text-to-SQL - Thinking Machines Lab
An expert-verified training set and a reshaped reward let RLVR on Tinker train a single model past human-level text-to-SQL accuracy, at a fraction of frontier cost.
Many industries rely on relational databases that are queried with SQL. Most SQL is machine-written, but humans alone likely write billions of custom SQL queries each monthBased on our internal estimates and publicly available data, such as from Snowflake filings. prompted by business questions. They are quite good at it — humans score 92.96% on BIRD, a realistic benchmark for translating natural-language questions into SQL. However, AI performance on text-to-SQL has lagged behind. LLM scores on the BIRD leaderboard improved from just below 70% in 2024 to 82% today. Frontier models like…
saved by
related reading
- BIRD-benchbird-bench.github.io
- As Rocks May Think | Eric Jangevjang.com
- Noisy Data Breaks RLVRddkang.substack.com
- Spider: Yale Semantic Parsing and Text-to-SQL Challengeyale-lily.github.io
- Composer2.pdfcursor.com
- PostTrainBenchposttrainbench.com
- Cookbookcookbook.openai.com
- defog/sqlcoder-7b · Hugging Facehuggingface.co
- DeepSeek-R1arxiv.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Expert Data for Frontier AI - AfterQueryafterquery.com
- LLMs shouldn’t write SQL - by Benn Stancil - benn.substackbenn.substack.com