Back

Under review · Under review

Isolating Architectural Effects in Text-to-SQL: A Controlled Diagnostic Study Using the Spider Sandbox

Under review (ARR), 2026


Abstract

Text-to-SQL evaluations routinely co-optimize model, prompt, retrieval, and generation strategy simultaneously, conflating fundamental model capability with system-level engineering. We evaluate three open-weight models (Gemma 4 31B dense, Qwen3-Coder 30B MoE with ~3B active, Qwen3-235B MoE with ~22B active) through an identical 11-step ADAPT-SQL pipeline used as a controlled diagnostic sandbox. A robust pipeline equalizes models on 89% of standard queries within 0.8pp despite a tenfold difference in active parameters. On nested-complex tasks, performance diverges by 4.4pp: MoE routing in Qwen3-235B activates richer reasoning circuits per token than a dense architecture of equivalent active scale, an ordering inconsistent with active-parameter count alone. We release the fixed-pipeline protocol as a reproducible framework for architectural ablations.

Text-to-SQLSpidermixture of expertscontrolled evaluationopen-weight LLMs