Query optimizers have been a known weak point in databases for a decade. Leis and colleagues asked how good optimizers really are in 2015, then asked again ten years later, and found the answer still disappointing. The hardest part, join ordering, is NP-hard. What is easy, though, is checking whether a chosen plan is good: run it and measure the execution time. That single verifiable axis makes query planning a natural fit for reinforcement learning. Rohan Bansal published an experiment on September 16 asking whether a small open-weights model can be post-trained to produce Postgres plans that beat the database's default planner.

The method is straightforward. His 4B Qwen model proposes a candidate strategy for each rollout, expressed as planner hints such as a hash join or a leading order. Postgres executes each candidate and measures it against its own default plan, and the resulting scalar rewards flow backwards to nudge the model's weights toward faster plans. Because a database benchmark is an inherently noisy reward signal, he designed a custom GRPO variant for scoring rollouts. Training was split across two machines: vLLM and the trainer on a rented two H100 node, and four Postgres containers running on his desk. He also ran off-policy distillation across roughly five hundred GPT-6 Astra agent trajectories.

The results are striking for a model this small. The tuned model cut latency by 44.7 percent across 113 join-heavy queries from the Join Order Benchmark and the Cardinality Estimation Benchmark, with individual plans running up to 81 percent faster than Postgres's default. The vanilla 4B baseline could not produce a usable plan for 99 of those 113 queries at all, so most of the gain came from simply teaching the model to plan before reinforcement learning made it fast. The conclusion is not that Postgres is obsolete, but that a narrow, measurable task plus a small open model plus verifiable rewards is a potent combination. Bansal's next steps point at applying the same loop to other database internals.