Skip to content

Fix SQLite-discovered SQL issues and support scalar subquery initialization - #389

Merged
KKould merged 16 commits into
mainfrom
fix/sqlite-slt-findings
Oct 8, 2026
Merged

KKould merged 16 commits into
mainfrom
fix/sqlite-slt-findings

Conversation

@KKould

@KKould KKould commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

What problem does this PR solve?

SQLite sqllogictest probing exposed incorrect index ranges, scalar-subquery limitations, and avoidable Cartesian products in implicit joins.

Issue link: Closes #386

What is changed and how it works?

  • Fix AND-over-OR index range intersections and NULL range ordering; display NULL as NULL.
  • Evaluate uncorrelated scalar subqueries once and correlated SELECT-list scalars per input row. Preserve row-specific values through memory sort, TopK, and spilled sorting.
  • Add callback-based Database::run_mut, reuse it for DDL, and remove SQL reprinting/reparsing from the SLT harness.
  • Add abs, eliminate unary + during binding, and use PostgreSQL-style integer division while preserving fractional AVG. Fix alias precedence, numeric casts, and boolean short-circuiting.
  • Recognize equijoins in filtered cross joins and increase predicate-pushdown iterations. Stream range combinations without intermediate nested vectors.

Code changes

  • Has Rust code change
  • Has CI related scripts change

Check List

Tests

  • Unit test
  • Integration test
  • Manual test (add detailed scripts or steps below)
  • No code

Side effects

  • Performance regression: Consumes more CPU
  • Performance regression: Consumes more Memory
  • Breaking backward compatibility

Note for reviewer

  • Validation: make test passes (495 library tests, 20 macro tests), and all 199 SLT files pass. External converted SQLite select1–3 pass 5,413 records; index samples pass 71,486, with 5 overflow failures and 10 recursion-limit skips.
  • Sorting retains per-row scalar values. Integer division, NULL display, evaluator numbering, and scalar-reference serialization change behavior/formats.
  • Correlated DISTINCT, outer aggregate/window scopes, and multi-level scalar correlation remain unsupported. Join reordering and external corpus tooling are not included; select4/5 still exceed bounded debug runs.

KKould added 11 commits October 8, 2026 04:21
An AND over a union was merged like an OR: pieces disjoint from the other
operand were kept, yet the predicate was reported fully consumed, so
EliminateIndexFilter dropped the filter and index scans returned extra rows
(e.g. id > 10 AND (id = 3 OR id = 21) returned 3). An empty piece in the
union also hit an unreachable!() arm.

Split range merging by operator: AND sweeps both sorted unions and keeps only
the intersections of overlapping pieces, OR keeps the in-place union merge.
NULL index keys sort after every value, but range merging disagreed:
- Eq x Eq placement used partial_cmp, so Eq(NULL) vs Eq(x) counted as
  overlapping; OR-merging them made no progress and recursed until the stack
  overflowed (e.g. c1 = 1 or c1 = 2 or c1 is null).
- Scope OR Eq(NULL) put NULL first, and kept it only when the lower bound was
  excluded, so c1 >= 3 or c1 is null and c1 < 3 or c1 is null dropped the
  NULL rows. A scope only covers NULL when its upper bound is unbounded.

Compare Eq x Eq with bound_compared, keep NULL as a trailing piece unless the
scope already covers it, and drop the special NULL arms in union_piece that
were masking the inconsistency.
Match SQLite, PostgreSQL and MySQL. Also show a missing column default as
NULL in describe, and update slt expectations (two rowsort blocks reorder
because NULL now sorts before lowercase text).
Replace ScalarApply and WHERE scalar joins with ScalarQueryInit and leaf Init expressions. Keep statement-local results in an owning ExecArenaView cache, including NULL results, and make them available before sorting, grouping and DISTINCT.

Use arena-scoped references for initialization definitions and relocate view references before serialization. Add subquery SLT regressions and a database-level catalog reload test. Retain explicit rejection of correlated scalar subqueries.
Register signed and unsigned integer and floating-point overloads. Preserve NULL values and reject signed integer absolute-value overflow. Cover numeric columns, arithmetic arguments, predicates, NULL and overflow in SLT.
Expose Database::run_mut to consume the final statement's results inside a callback, reusing DDL execution, commit and catalog publication. Move statement scope cleanup into TransactionGuard so arena and DDL changes can be returned by value.

Use run_mut in the SLT harness to avoid SQL parse-display-parse roundtrips. Reject aggregates in WHERE, prefer input columns over conflicting SELECT aliases in SELECT/WHERE, and correct floating-point integer-cast range checks with regression coverage.
Keep integer division results integral with checked zero and overflow handling, while preserving fractional AVG and floating-point results. Eliminate unary plus during binding and remove its evaluator dispatch. Short-circuit decisive boolean operands so guarded division is not evaluated.

Update division expectations and cover text identity, signed division, overflow, NULL propagation and mixed floating-point expressions.
Push single-input predicates through cross joins and extract conjunctive equality predicates as inner-join keys, retaining residual filters and localizing right-side positions. Share join-key extraction with explicit JOIN binding.

Raise the predicate-pushdown iteration cap to 30 for deeper join trees. Cover plans and query results with join SLT regressions. This does not implement join-order optimization.
Capture outer parameters separately from scalar results and use OuterParam, OuterValue and InitValue expressions to read execution-local values. Attach scalar initializers before their consumers without tuple result-column placeholders or post-hoc projections.

Use an explicit ExecMetaArena on execution paths and confine value updates to it. Model scalar initialization with states and reuse a scratch tuple buffer while running child plans. Retain explicit rejection of correlated ORDER BY, DISTINCT and unsupported outer scopes.

Add subquery SLT coverage for stage attachment, missing/NULL/text results, multiple correlations and unsupported boundaries. Gate the LMDB bounds helper to its supported build configurations.
Mark row-scoped scalar references and move their values alongside cached sort rows, restoring them before projection. Cover memory sort, TopK and spilled merge runs without scanning plans or introducing separate value maps.

Flatten memory-sort keys, compose spill codecs for references and pairs, and consume range combinations directly through callbacks. Extend scalar-subquery regressions including NULL, text and mixed constant/row-dependent results.
@codecov

codecov Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.44423% with 134 lines in your changes missing coverage. Please review.
✅ Project coverage is 93.68%. Comparing base (65d9638) to head (0d528f8).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
src/planner/arena.rs 83.63% 36 Missing ⚠️
src/expression/eq_col.rs 38.70% 19 Missing ⚠️
src/binder/select.rs 93.58% 12 Missing ⚠️
src/expression/range_detacher.rs 94.50% 11 Missing ⚠️
src/planner/operator/join.rs 87.91% 11 Missing ⚠️
src/binder/expr.rs 90.36% 8 Missing ⚠️
src/optimizer/rule/normalization/column_pruning.rs 86.95% 6 Missing ⚠️
src/execution/spill/codec.rs 90.90% 5 Missing ⚠️
...ptimizer/rule/normalization/pushdown_predicates.rs 92.06% 5 Missing ⚠️
src/function/abs.rs 91.83% 4 Missing ⚠️
... and 8 more
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #389      +/-   ##
==========================================
+ Coverage   93.30%   93.68%   +0.38%     
==========================================
  Files         259      260       +1     
  Lines       47765    48783    +1018     
==========================================
+ Hits        44565    45704    +1139     
+ Misses       3200     3079     -121     
Flag Coverage Δ
rust 93.68% <93.44%> (+0.38%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
src/binder/aggregate.rs 98.00% <100.00%> (+0.32%) ⬆️
src/binder/create_view.rs 100.00% <100.00%> (ø)
src/binder/mod.rs 86.74% <100.00%> (+0.18%) ⬆️
src/catalog/column.rs 90.85% <100.00%> (-1.83%) ⬇️
src/db.rs 93.42% <100.00%> (+1.38%) ⬆️
src/db/prepared.rs 94.55% <100.00%> (ø)
src/execution/ddl/add_column.rs 96.93% <100.00%> (ø)
src/execution/ddl/change_column.rs 97.08% <100.00%> (ø)
src/execution/ddl/create_index.rs 91.25% <100.00%> (ø)
src/execution/ddl/create_table.rs 100.00% <100.00%> (ø)
... and 83 more

... and 9 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

KKould added 5 commits October 8, 2026 04:31
Keep shallow CTE plan references instead of cloning every expression at bind time. Remap positions through consumer-local expression references and reuse rewrites within each input schema. Cover repeated CTEs whose consumers require different columns.
Clone the CTE plan structure once and recursively isolate its mutable expressions using ExprCloner. Preserve scalar initialization references and cover repeated CTE consumers with different pruning requirements.
@KKould KKould self-assigned this Oct 8, 2026
@KKould KKould added the enhancement New feature or request label Oct 8, 2026
@KKould
KKould merged commit c78efe9 into main Oct 8, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Evaluate uncorrelated scalar subqueries once as constants

1 participant