arXiv cs.CL
7/1/2026

Test-Time Verification for Text-to-SQL via Outcome Reward Models
Short summary
Researchers introduce GradeSQL, a framework using Outcome Reward Models (ORMs) to verify Text-to-SQL outputs at inference time, replacing heuristic selection strategies. On BIRD and Spider benchmarks, ORM-based selection outperforms execution-based approaches by up to 4.33%, with stronger gains on complex queries. Code and datasets are publicly available.
- •GradeSQL uses learned semantic scoring (ORMs) instead of heuristic signals for Text-to-SQL verification
- •Achieves +4.33% on BIRD and +2.10% on Spider benchmarks over baseline approaches
- •Scales effectively with larger candidate sets and includes public code/datasets
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


