Conversation
…ery NULLs `<constant> NOT IN (<subquery>)` in a `WHERE` clause returned every row when the subquery produced a NULL. `3 NOT IN (1, NULL)` is UNKNOWN, so the clause must remove the row; DuckDB and PostgreSQL return no rows. The value expression `3` has no column reference, so `find_valid_equijoin_key_pair` rejects `Int64(3) = __correlated_sq_1.id` and it stays in the join filter. Two things then went wrong: * The filter references only the subquery side, so `push_down_filter` moved it into the subquery as `Filter: t2.id = 3`, dropping the NULL rows before the join could observe them. * The join was left without equi-join keys and was planned as a `NestedLoopJoinExec`, which has no null-aware implementation, so the `null_aware` flag was silently discarded. Three changes: * `DecorrelatePredicateSubquery` projects a constant value expression as a column of the outer input, so the predicate becomes a real equi-join key and the existing null-aware hash join handles it. The rewrite is limited to uncorrelated subqueries: a correlation predicate would be a second join key and null-aware hash joins accept only one. * `push_down_all_join` never pushes predicates into the right input of a null-aware join, matching the check `infer_join_predicates` already has. * The physical planner returns an error instead of building a keyless null-aware join that silently ignores the flag. Regression tests cover the anti-join and mark-join shapes, the controls from the report, a user column colliding with the projected value column, and the remaining unsupported case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VCeMPyJNgAiFpGXaz5CxF3
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
A follow-up issue covers the remaining
NOT INproblems. Its text is ready, but this repository has issues disabled, so it is not filed yet. See Follow-up work at the end.Rationale for this change
3 NOT IN (1, NULL)is UNKNOWN. AWHEREclause must remove the row. DataFusion kept every row. There was no error and no warning.1, 2❌The subquery gives
{1, NULL}. The value3is notNULL, but3is also not known to be absent. Therefore the answer is UNKNOWN for every row oft1.The same expression in a
SELECTlist was already correct. Only theWHEREclause was wrong.Why it happened
The value
3holds no column. Therefore it could not become a join key. It stayed as a join filter. Two problems followed.What changes are included in this PR?
Three changes.
1. Make the constant a join key (
datafusion/optimizer/src/decorrelate_predicate_subquery.rs)Add the constant to the outer side as a column. The comparison then becomes a true equality of two columns, so the existing null-aware hash join does the work.
This applies only to uncorrelated subqueries. A correlated subquery needs a second key, and a null-aware anti join accepts only one key.
2. Keep the filter out of the subquery (
datafusion/optimizer/src/push_down_filter.rs)Do not move a predicate into the subquery side of a null-aware join. The NULLs must reach the join.
infer_join_predicateshas the same rule already.3. Report an error instead of a wrong answer (
datafusion/core/src/physical_planner.rs)Only a hash join can do null-aware work, and it needs a key. If a null-aware join has no key, report an error. Do not build a nested loop join that gives wrong results without a warning.
What is the testing strategy for this PR?
New
sqllogictestcases indatafusion/sqllogictest/test_files/null_aware_anti_join.sltandnull_aware_mark_join.slt. They cover:ORandIS NULLforms, which use a mark join;New unit tests in
decorrelate_predicate_subquery.rsandpush_down_filter.rshold the new plans.Results:
sqllogictestsuitedatafusion-optimizerunit testscargo fmt --allcargo clippy --all-targets --all-features -- -D warningsdatafusion-substraitis not in the extended run, becauseprotocis absent from the build machine. This PR changes no serialization code.Are there any user-facing changes?
Yes. Two.
1.
<constant> NOT IN (<subquery>)in aWHEREclause now gives correct results. This is the fix.2. One shape now reports an error. A constant with a correlation that is not an equality:
This query gave wrong results before. No correct result is lost. An error is better than a wrong answer that a user cannot see.
There are no API changes.
Follow-up work
This PR makes every uncorrelated
NOT INcorrect. CorrelatedNOT INstill has problems. They come from the join operator, not from the code this PR changes.A test matrix of 18 query shapes gives these totals:
The 10 remaining shapes are all correlated. They fall into three parts:
NOT INneeds two.Part 3 must come first, as an error. A prototype showed that part 2 alone turns a clear error into a silent wrong answer.
🤖 Generated with Claude Code
https://claude.ai/code/session_01VCeMPyJNgAiFpGXaz5CxF3
Generated by Claude Code