Skip to content

fix(sql): an empty group's aggregate is NULL in HAVING; MATCH in HAVING is refused by name - #408

Merged
fupelaqu merged 1 commit into
mainfrom
fix/having-null-and-match
Oct 1, 2026
Merged

fupelaqu merged 1 commit into
mainfrom
fix/having-null-and-match

Conversation

@fupelaqu

@fupelaqu fupelaqu commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

What

An empty group's aggregate is NULL in HAVING. For a group in which no document has the column, Elasticsearch hands the group filter (bucket_selector) NaN for its MIN, MAX, AVG or percentile, not NULL. So <>, a negated comparison, IS [NOT] NULL and COALESCE kept or dropped the wrong groups, with HTTP 200 and no error.

  • The filter now reads null, NaN and ±Infinity as NULL and applies SQL three-valued logic: a comparison with NULL is UNKNOWN, NOT UNKNOWN stays UNKNOWN, and a group is kept only when the condition is TRUE. This is exact for indexed data: Elasticsearch 6.8–9.0 index only finite numbers.
  • A NULL-handling function is checked on its result. COALESCE(MAX(a), MIN(b)) is NULL only when both are.
  • No aggregation is added. COUNT and SUM are read as before (0 and 0.0).
  • Search queries (both bridges) and materialized views (Having.script) build the filter through one function, MetricSelectorScript.nullAwareSelectorScript, and metricSelector delegates to it.

MATCH in HAVING is refused by name. A MATCH in HAVING over an aggregate, or over a column that is neither an aggregate nor a GROUP BY key, was silently ignored. It is now refused with the remedy "Put the MATCH in WHERE." A MATCH over a GROUP BY key is unchanged.

Docs: documentation/sql/dql_statements.md (the HAVING rules and a "Changed in 0.24.0" note).

Measured (ES 8.18.3 through the gateway unless stated)

Check main this branch
NULL population: 1,560 statements over every aggregate that is NULL on an empty group (groups with the column absent, partly present, fully present; independent three-valued oracle) 425 statements / 487 groups wrong 0 wrong, 0 errors
COALESCE population: 398 statements 300 statements / 464 groups wrong 0 wrong, 0 errors
New refusals across 3,590 parsed statements — only the 12 MATCH statements; on main each answer equalled the answer without the MATCH (12/12)
Materialized-view filter vs search filter — identical on every statement; no new buckets_path variable
WHERE output (401 statements, 244 scripts) — identical to main
ES execution, 200 HAVING bodies, 1M docs, force-merged, request_cache=false, interleaved — summed took medians ×0.998; none slower beyond noise
Emission probe, 1,628 HAVING statements, 3 interleaved rounds against a control — ×1.03, ranges overlap (×1.005 in an earlier round); parser unchanged

Suites: sql 1765, es6bridge 256, softclient4es8-sql-bridge 263, core 1148. ES 8.18.3: the population spec (53 tests) and 5 HAVING integration suites (261). Lint (headerCheck scalafmtSbtCheck scalafmtCheck test:scalafmtCheck) clean.

Not run locally, left to CI: ES 6.8 / 7.17 / 9.0 (the population spec runs on every major), the Scala 2.12 cross-compile, and the es7 / es9 bridges.

Behaviour changes (release notes)

  • HAVING over an empty group's MIN / MAX / AVG / percentile follows SQL NULL semantics. Statements using <>, a negated comparison, IS [NOT] NULL or COALESCE may return other groups than 0.23.0, which returned wrong ones.
  • New refusal: MATCH in HAVING over an aggregate, or over a column that is neither an aggregate nor a GROUP BY key. Remedy: put it in WHERE.
  • Materialized views: the generated HAVING filter script changes for these aggregates. A view picks it up when it is (re)deployed.
  • API: MetricSelectorScript.metricSelector now returns the NULL-aware script, and Having.script's deprecation points to MetricSelectorScript.nullAwareSelectorScript. No case-class arity change.

Other pre-existing HAVING limitations found during this work are unchanged here and kept for later.

🤖 Generated with Claude Code

…NG is refused by name

For a group in which no document has the column, Elasticsearch hands the
HAVING filter NaN for its MIN, MAX, AVG or percentile. The filter now
reads null, NaN and +/-Infinity as NULL and follows SQL three-valued
logic, so <>, a negated comparison, IS [NOT] NULL and COALESCE keep the
groups SQL keeps. A NULL-handling function is checked on its result. No
aggregation is added. Search and materialized views build the filter
through one function, nullAwareSelectorScript, and metricSelector
delegates to it.

A MATCH in HAVING over an aggregate, or over a column that is neither an
aggregate nor a GROUP BY key, was silently ignored. It is now refused by
name, with the remedy to put it in WHERE.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@fupelaqu
fupelaqu marked this pull request as ready for review October 1, 2026 19:18
@fupelaqu
fupelaqu merged commit 47417b3 into main Oct 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant