Conversation
Bumps [cxx](https://github.com/dtolnay/cxx) from 1.0.200 to 1.0.202. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/dtolnay/cxx/releases">cxx's releases</a>.</em></p> <blockquote> <h2>1.0.202</h2> <ul> <li>Support build environments that limit build-script filesystem writes outside of OUT_DIR (<a href="https://redirect.github.com/dtolnay/cxx/issues/1760">#1760</a>, thanks <a href="https://github.com/gudcks0305"><code>@gudcks0305</code></a>)</li> </ul> <h2>1.0.201</h2> <ul> <li>Specialize <code>enable_borrowed_range</code> and <code>enable_view</code> for <code>rust::Slice<T></code> (<a href="https://redirect.github.com/dtolnay/cxx/issues/1758">#1758</a>, thanks <a href="https://github.com/MixusMinimax"><code>@MixusMinimax</code></a>)</li> <li>Register bazel toolchains as dev dependencies (<a href="https://redirect.github.com/dtolnay/cxx/issues/1759">#1759</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/dtolnay/cxx/commit/983982c5fd5584a5391ae46877e952118e431f73"><code>983982c</code></a> Release 1.0.202</li> <li><a href="https://github.com/dtolnay/cxx/commit/552330e2a705f1b0ed03454eb2fc9b4d702ff556"><code>552330e</code></a> Touch up PR 1760</li> <li><a href="https://github.com/dtolnay/cxx/commit/3a9d010450c5ee145a9394ced0b83a294215a91e"><code>3a9d010</code></a> Merge pull request 1760 from gudcks0305/codex/fix-shared-header-fallback</li> <li><a href="https://github.com/dtolnay/cxx/commit/97c1cfffcbc51bea167b927dd4acdba147ef3bbf"><code>97c1cff</code></a> fix: tolerate unavailable shared header directory</li> <li><a href="https://github.com/dtolnay/cxx/commit/982aa18796aef5ffea9c26466bd21d8397002cb9"><code>982aa18</code></a> Release 1.0.201</li> <li><a href="https://github.com/dtolnay/cxx/commit/a543f3b5c859b6a11757a4899a0dd367fb2a7a37"><code>a543f3b</code></a> Lockfile update</li> <li><a href="https://github.com/dtolnay/cxx/commit/a237c373198ff5cfe4aa8dc9e2e49ac75643b114"><code>a237c37</code></a> Reformat with clang-format 22</li> <li><a href="https://github.com/dtolnay/cxx/commit/eeabfb536ec143ff52bd1c54a2bc9eecba6881f2"><code>eeabfb5</code></a> Move static assertions from header to source file</li> <li><a href="https://github.com/dtolnay/cxx/commit/47d9bf7a5fae6657712e1e1373f1758b40fca660"><code>47d9bf7</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/cxx/issues/1758">#1758</a> from MixusMinimax/slice-impl-borrowed-range-and-view</li> <li><a href="https://github.com/dtolnay/cxx/commit/5a9c2ba4cebbc3ac62f20175e1bb60c26b82affb"><code>5a9c2ba</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/cxx/issues/1759">#1759</a> from dtolnay/bazel</li> <li>Additional commits viewable in <a href="https://github.com/dtolnay/cxx/compare/1.0.200...1.0.202">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…cks (sirius-db#1800) Bumps [cxx](https://github.com/dtolnay/cxx) from 1.0.200 to 1.0.202. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/dtolnay/cxx/releases">cxx's releases</a>.</em></p> <blockquote> <h2>1.0.202</h2> <ul> <li>Support build environments that limit build-script filesystem writes outside of OUT_DIR (<a href="https://redirect.github.com/dtolnay/cxx/issues/1760">#1760</a>, thanks <a href="https://github.com/gudcks0305"><code>@gudcks0305</code></a>)</li> </ul> <h2>1.0.201</h2> <ul> <li>Specialize <code>enable_borrowed_range</code> and <code>enable_view</code> for <code>rust::Slice<T></code> (<a href="https://redirect.github.com/dtolnay/cxx/issues/1758">#1758</a>, thanks <a href="https://github.com/MixusMinimax"><code>@MixusMinimax</code></a>)</li> <li>Register bazel toolchains as dev dependencies (<a href="https://redirect.github.com/dtolnay/cxx/issues/1759">#1759</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/dtolnay/cxx/commit/983982c5fd5584a5391ae46877e952118e431f73"><code>983982c</code></a> Release 1.0.202</li> <li><a href="https://github.com/dtolnay/cxx/commit/552330e2a705f1b0ed03454eb2fc9b4d702ff556"><code>552330e</code></a> Touch up PR 1760</li> <li><a href="https://github.com/dtolnay/cxx/commit/3a9d010450c5ee145a9394ced0b83a294215a91e"><code>3a9d010</code></a> Merge pull request 1760 from gudcks0305/codex/fix-shared-header-fallback</li> <li><a href="https://github.com/dtolnay/cxx/commit/97c1cfffcbc51bea167b927dd4acdba147ef3bbf"><code>97c1cff</code></a> fix: tolerate unavailable shared header directory</li> <li><a href="https://github.com/dtolnay/cxx/commit/982aa18796aef5ffea9c26466bd21d8397002cb9"><code>982aa18</code></a> Release 1.0.201</li> <li><a href="https://github.com/dtolnay/cxx/commit/a543f3b5c859b6a11757a4899a0dd367fb2a7a37"><code>a543f3b</code></a> Lockfile update</li> <li><a href="https://github.com/dtolnay/cxx/commit/a237c373198ff5cfe4aa8dc9e2e49ac75643b114"><code>a237c37</code></a> Reformat with clang-format 22</li> <li><a href="https://github.com/dtolnay/cxx/commit/eeabfb536ec143ff52bd1c54a2bc9eecba6881f2"><code>eeabfb5</code></a> Move static assertions from header to source file</li> <li><a href="https://github.com/dtolnay/cxx/commit/47d9bf7a5fae6657712e1e1373f1758b40fca660"><code>47d9bf7</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/cxx/issues/1758">#1758</a> from MixusMinimax/slice-impl-borrowed-range-and-view</li> <li><a href="https://github.com/dtolnay/cxx/commit/5a9c2ba4cebbc3ac62f20175e1bb60c26b82affb"><code>5a9c2ba</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/cxx/issues/1759">#1759</a> from dtolnay/bazel</li> <li>Additional commits viewable in <a href="https://github.com/dtolnay/cxx/compare/1.0.200...1.0.202">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…ge extraction (sirius-db#1577) # perf(scan): see through DATE→TIMESTAMP casts in scan filter range extraction **Branch:** `pr/cast-through-range-extraction` → base: **`dev`**. Two commits, no longer stacked: the change, and the review follow-up for @mbrobbel's ±infinity / overflow comment. Rebased onto `dev` @ `05148c78` (2026-09-15). > [!NOTE] > **Previously stacked on sirius-db#1555**, which has since merged (`3ac080d8`). This PR was rebased onto `dev`: the ~70 never-merged commits of the sirius-db#1391 fused scan-filter lineage are gone, and the one remaining change was **ported** rather than replayed — `dev` had meanwhile refactored the extraction machinery out of `scan_utils.cpp` into `op/scan/scan_filter_analysis.{hpp,cpp}` (`extract_numeric_range_pushdown` → `analyze_scan_filters`, `decoded_bound{floor,ceil}` → `sirius::numeric_range{minimum,maximum}`). The diff below is now exactly this change: **5 files, +889 / −24**. ## Problem DuckDB constant-folds qgen-style date arithmetic (`l_shipdate <= DATE '1998-12-01' - INTERVAL '72' DAY`) into `CAST(l_shipdate AS TIMESTAMP) <= TIMESTAMP '1998-09-20 00:00:00'`, which arrives at the scan as an `EXPRESSION_FILTER`. `analyze_scan_filters` only understood `CONSTANT_COMPARISON` / `CONJUNCTION_AND`, so `fold_numeric_conjunct` refused the clause and the whole scan lost its in-decode masking: full-width decode + per-row cast AST filter + gather-compact. q1 varied 1.288 s vs 0.831 s with the equivalent DATE literal; the same shape appears in 9 predicates across q1/q4/q5/q6/q10/q12/q14/q15/q20 under spec-official (qgen) parameters. ## Change Three additions in `src/op/scan/scan_filter_analysis.cpp`: - **`to_decoded_bound`**: accepts finite `TIMESTAMP[_S|_MS|_NS]` constants against DATE columns, lowering ticks / ticks-per-day into the stored-day domain by floor/ceil (templated on the finer-scale DECIMAL branch). The bound is EXACT for every comparison op — midnight(d) is strictly monotonic in d — so midnight and non-midnight constants both keep full coverage and the residual filter drops; `= non-midnight` folds to a provably empty range. - **`fold_comparison_bound`**: the comparison switch lifted out of `fold_numeric_conjunct` so both filter shapes share one implementation. - **`fold_expression_conjunct`** + a new `EXPRESSION_FILTER` case, matching `cmp(CAST(bound_ref AS TIMESTAMP*), ts_const)` in both operand orders (comparison flipped on swap), recursing through AND conjunctions. **Edges of the DATE domain** (review follow-up, see `lower_timestamp_to_days`): `DATE ±infinity` rows compare correctly because ±INT32_MAX days and ±INT64_MAX ticks are the extremes of their domains, and **±infinity constants now lower exactly** to the infinite date's stored days instead of being refused. A finite date whose midnight overflows the target (~±292,000 years for `TIMESTAMP/_S/_MS`, ~±292 years for `TIMESTAMP_NS`) makes DuckDB raise; a decode range cannot raise, and neither does the residual `cudf::cast` it replaces, so the range is deliberately not clipped to the castable window and such rows are kept or dropped by the instant they denote. The only divergence from DuckDB, and only on queries DuckDB refuses to answer. **Refused (coverage cleared, residual kept):** `TRY_CAST` (NULL-on-overflow diverges from day math on >-side bounds), `TIMESTAMP_TZ` (midnight is session-timezone dependent), `INT64_MIN` ticks (below -infinity, never a real instant), and every non-comparison shape (including the `BETWEEN` DuckDB builds from two cast comparisons on one column; follow-up). ## Measured (SF1000, GB300, spec-official qgen parameters) - **Power@Size +9.6% → 10,079,305** - Power-run suite time **−8.3%** - q1/q12/q14/q15 each **−36..39%** (the queries whose only selective predicate is a folded date cast) > [!WARNING] > These numbers were measured on the **pre-rebase fused-scan-filter stack**, not on this branch. They have **not** been re-measured against current `dev`, whose scan/decode path has since changed (sirius-db#1524 late materialization, sirius-db#1555, and others). Treat them as the original motivation, not as a claim about this branch's delta. Correctness on current `dev` is verified below. ## Amplification history (honest note) Multiplying the number of fused scans exposed two classes of pre-existing bugs downstream of scan-produced geometry — neither introduced by this change, both diagnosed during its qualification: - two pre-existing Sirius GPU use-after-free races (batch read-lock lifetime, error-path teardown); - a pre-existing **CCCL 3.4.0** bug: the warp-specialized (TMA) `cub::DeviceScan` on Blackwell performs an out-of-bounds `__shared__` read that silently corrupts scan outputs under concurrent kernel diversity (memcheck-reproducible with a 22-line bare `ExclusiveSum`; CTK-bundled CCCL 3.2.0 is clean). Both hardenings ship separately in **sirius-db#1580** (which supersedes the now-closed sirius-db#1546). ## Tests - `test/cpp/scan/test_scan_filter_cast_ranges.cpp` (`[range_pushdown]`): exhaustive boundary matrix — all 5 ops × midnight / one-µs-either-side / noon constants, pre-epoch floor, operand flips, AND shapes, S/MS/NS flavors, extreme finite ticks, ±infinity constants (all ops, both directions, both operand orders, every flavor), `INT64_MIN` refusal, bounds not clipped to the castable window, and the refusal set. - `test/cpp/integration/test_gpu_execution_cast_date_predicates.cpp` (`[cast_date]`): GPU-vs-CPU equivalence for the folded SQL shapes on three scan paths. **On current `dev` the fold's ranges only take effect on a GPU-pinned Parquet entry decoded through the compressed path** (the DuckDB-native ingestible does not analyze its filter), so besides the plain and DuckDB-pinned scans the fixture pins bitpacked Parquet twins, proves via census that the pin compressed, and lifts the fused-decode selectivity caps so the in-decode mask answers. New tables: `DATE ±infinity` rows (finite constants on every path, infinite constants on the fold path) and a year-300000 date that raises on CPU while the fold answers by instant. Run on this branch, rebased onto `dev` @ `05148c78`: | Filter | Result | | --- | --- | | `[range_pushdown],[cast_date],[fused_scan_filter]` | 17/17 green — 9,788 assertions | | `[scan],[filter]` neighbor sweep | 374/374 green — 461,053 assertions | `pixi run pre-commit run -a` clean on the changed files. > [!NOTE] > Found while making the tests bite, pre-existing and not touched here: the residual `cudf::cast` path multiplies `±INT32_MAX` days into int64 micros and wraps, so e.g. `d < TIMESTAMP 'infinity'` returns the +infinity row on unpinned / DuckDB-format scans today. The fold gives the right answer where it engages. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Handle breaking changes in 26.12 to fix nightly job.
…sion exploration (sirius-db#1726) ## Description This PR brings in the changes from Felipe's late-mat branch sirius-db#1391 to fix and improve simpatico's explore function. It includes excluding JIT costs from the evaluation, and outputting all pareto-optimal points so we can change the throughput floor without re-measuring. ## Checklist - [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist - [x] Cover changes with new or existing tests - [x] Document configuration changes in code and summarize in the description above - [x] Update human and agent documentation (README.md, docs/, skills, CLAUDE.md) ## References Co-authored-by: Felipe Aramburu <faramburu@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…up metadata walk for pin-served scans (sirius-db#1548) ## Description Two independent frontend fixes, one per commit, that cut redundant plan-construction work out of the query path. ### 1. Build the Sirius physical plan once per query The transparent path built the Sirius GPU physical plan **twice per query**: once in `OnFinalizePrepare` purely to validate GPU support (the resulting plan was discarded), and again in `PhysicalSiriusExecution::GetDataInternal` for the actual run. The validated plan is now handed to `PhysicalSiriusExecution` and consumed by the first `GetData` (one-shot — the move empties the slot). Re-executions of the same prepared operator rebuild from the logical plan template exactly as before, so data-drift and rebind semantics are unchanged. ### 2. Defer the row-group metadata walk for pin-served scans Every plan build constructs one duckdb-native ingestible per `seq_scan`, whose constructor ran `prepare_duckdb_native_walk`: `GetPartitionStats` plus a per-row-group statistics pass per prunable filter column and per projected varchar column. For a **pin-served scan its output is never consumed** — splits come from the pinned entry's chunk plan and insert-delta bundles, not from the walk. `sirius_plan_get` now marks a scan whose MVCC cache-or-CPU guards all passed with `mvcc_pin_serves_scan`; the plan generator forwards it as `defer_metadata_walk`; the ingestible skips the walk at construction; and the new `gpu_ingestible::ensure_metadata_prepared()` hook runs it lazily from the scan manager, on the query thread, for exactly the scans that did *not* match a pinned entry at execution time. A scan that actually reads disk always gets its walk. The two compound: fix 1 halves the number of plan builds, fix 2 removes the walk from the remaining build on pinned workloads. ### Refusal semantics preserved - The projected-type gate stays **eager in both modes** (factored out as `unsupported_projected_type_reason`), so unsupported types still refuse at plan time, where the throw is a clean CPU fallback. - `sirius_plan_get`'s table-level overflow-string probe strictly dominates the walk's per-row-group overflow check. For deferred scans the per-row-group refusal moves to the lazy walk (a runtime scan error → CPU fallback); for unpinned scans it stays at construction. - A lazy-walk failure throws out of `prepare_for_query` and surfaces as a runtime scan error, matching the walk's other runtime refusals. No new configuration knobs; behavior is unconditional. ## Measurements **Re-measured 2026-09-16 after rebasing onto `dev` @ `05148c78`** (RAPIDS 26.08), same GB300 box, same duckdb-native SF1000 database, unpinned, q1–q19, best-of-3 per query, arms alternated, each arm built in its own worktree. Only runs where the process was in the box's fast mode are comparable (see the caveat below): | | dev @ 05148c7 | this branch | | |---|---|---|---| | suite total, run A | 36.362 s | 35.907 s | | | suite total, run B | 36.696 s | 35.922 s | | | suite total, min | 36.362 s | 35.907 s | **−1.25%** | | queries equal or faster | | | 19/19 | Largest movers: q12 −4.2%, q6 −3.6%, q5 −2.2%, q3 −1.9%. No query slower outside its own run spread. This matches the −1.05…−1.23% measured on the previous base — as expected, fix 1 is the whole effect on an unpinned workload. **Caveat — bimodal box, not bimodal code.** During this session the machine put roughly a third of launched Sirius processes into a persistent "slow mode": cold first iterations identical, every warm iteration ~2.6× slower, for the process's whole lifetime, on *either* binary. Twelve alternating single-process probes (q6+q12, 3 iterations) split dev 2 fast / 4 slow and this branch 4 fast / 2 slow, with the two binaries equal inside each mode (q6 1.66–1.71 s vs 1.66–1.71 s fast; 4.4–4.5 s vs 4.4–4.5 s slow). Slow-mode runs are excluded from the table above. The mechanism is not diagnosed here (best guess: host-pool placement vs page cache); it predates and is independent of this PR. **Also on both arms:** the pre-existing `SIGSEGV` in `cuMemcpyBatchAsync_v2` under `decode_duckdb_native_split` now also fires intermittently on q2 (and q9), not only q20/q3. Dev bug, reproduces on the base binary. <details><summary>Original measurements against dev @ 9624f06</summary> Measured on a GB300 workstation against a duckdb-native TPC-H SF1000 database, comparing this branch against its base (`dev` @ `9624f063`), each arm built in its own worktree. **Unpinned — the case where the walk cannot be skipped and must run anyway.** Best-of-3 per query, 4 complete runs on the base and 3 on this branch, arms alternated: | | base | this branch | | |---|---|---|---| | suite total, min of runs | 37.784 s | 37.389 s | **−1.05%** | | suite total, median of runs | 37.882 s | 37.416 s | **−1.23%** | | per-run spread | 37.78–38.15 s | 37.39–37.46 s | non-overlapping | | queries faster | | | 15/19 | The two arms' spreads do not overlap, so this is not run-to-run noise. Largest movers: q6 −4.1%, q12 −3.4%, q19 −2.3%, q4 −2.2%. Two queries moved the other way (q14 +3.2%, q13 +2.5%) but both sit inside their own per-run spread. This is fix 1 acting alone: with nothing pinned, `defer_metadata_walk` is never set and the walk runs eagerly on both arms — but the base runs it *twice per scan* and this branch once. Counting the per-operator plan-build log marker on a single-scan query confirms the mechanism: **6 emissions on the base, 3 on this branch** (3 logical operators × 2 builds vs × 1). **Pinned.** Verified by log markers that the feature engages as designed — one `deferring metadata walk` and one `reusing finalize-validated Sirius plan` per query, with the walk skipped entirely. The magnitude of the pinned-workload win is **not** re-measured here; the numbers previously quoted on this PR were taken against a superseded base and have been removed rather than carried forward. **Not measured:** the deferred-walk-then-pin-missing path (the walk runs from `prepare_for_query` instead of at plan time). Reaching it requires a pin to disappear between plan and execute, which I could not force deterministically at scale. It runs the same function on the same thread, just later. **Unrelated pre-existing failure:** q20 segfaults on the unpinned duckdb-native path on *both* arms (4/4 runs), and q3 does so intermittently — `SIGSEGV` in `cuMemcpyBatchAsync_v2` under `decode_duckdb_native_split`. It reproduces identically on `dev`, so it is not from this PR; it is why the table above covers q1–q19. </details> ## Checklist - [x] Read CONTRIBUTING.md - [x] Cover changes with new or existing tests - [x] Document configuration changes in code and summarize in the description above - [ ] Update human and agent documentation (README.md, docs/, skills, CLAUDE.md) — no user-facing behavior or configuration changes to document ## Tests - New `test/cpp/scan/test_duckdb_native_deferred_walk.cpp` (`[duckdb_native_deferred_walk]`, CPU-side, file-backed duckdb): deferred construction skips the walk; `ensure_metadata_prepared()` produces a plan identical to the eager one (row groups, starts, counts, filter-stat pruning); unsupported types still refuse at construction in both modes; overflow-string refusal placement per mode; a failed walk is retried rather than latched; deferred split claims match eager claims. - **AC-12 lifecycle-test adaptation** (`test/cpp/integration/test_query_lifecycle_slot.cpp`): the runtime-window parser required a "Creating sirius physical plan" line inside every execution window. With plan reuse the first execution's plan is created at finalize, so the parser now also accepts the plan-reuse line as plan provenance. The acceptance criterion itself is unchanged — the id-zero check and the cross-execution operator-id-sequence-equality check still verify that operator ids restart for every repeated execution. - Local: `[duckdb_native_deferred_walk]` + `[duckdb_native_walker]` 10 cases / 101 assertions pass; full default suite green (2828 cases, 32.7M assertions). ## References --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…#1767) This is the first step towards building Sirius as standalone library. Closes sirius-db#1732.
…capture (sirius-db#1804) ## Description Sirius emitted every NVTX range into the default domain, so profiles interleaved Sirius with `libcudf`/`CCCL` and couldn't be filtered to Sirius alone. **NVTX `sirius` domain** - All 80 annotation sites move into a named `sirius` domain. - The pipeline range spans task boundaries and can't be scoped, so it switches to `nvtxDomainRangeStartEx`/`End` — semantics unchanged. - `src/legacy/` is deliberately left on the global domain. **The domain tag type is declared twice on purpose** - simpatico also builds standalone, without `src/include` on its include path. - `parquet_benchmark` gets `src/include` but not simpatico's include root. - So neither header can include the other — collapsing them fails with `fatal error: codegen/util/nvtx.hpp: No such file or directory`. - NVTX keys domains by name, so both tag types resolve to one `sirius` domain. - The name comes from a single `SIRIUS_NVTX_DOMAIN_NAME` definition set from `${PROJECT_NAME}`, so it can't drift. **quent bump** - `381e1be` → `b21933d0`, for [quent#736](rapidsai/quent#736), which groups explicitly named NVTX domains by name at the presentation boundary. **Quent context ordering** - cuCascade's ranges gained a `libcucascade` domain in [cuCascade#196](NVIDIA/cuCascade#196), already on `dev` — but that domain never reached Quent. - `nvtx_injection::dispatch` drops events while its `OnceLock` hook is unset, and that hook is installed by `quent::create_context`, which ran inside `telemetry_context`'s constructor — i.e. *after* the memory manager was built. cuCascade's pinned-pool setup, including the `nvtxDomainCreateA` that names the domain, landed in that dead window. - The handle stayed valid, so later cuCascade ranges were captured but unnamed; the UI showed them under `domain0x1`. - `telemetry_context::create` now takes an already-built `rust::Box<quent::Context>`, and `SiriusContext::initialize` builds it first via the new `make_quent_context()`. This keeps the memory manager and per-GPU device ids the context declaration needs, so no telemetry is lost. **Verification** - TPC-H SF100 (Q1/Q3/Q5) under nsys: one `sirius` domain alongside `libcucascade`, `libcudf` and `CCCL`; zero ranges left in the default domain. - Under Quent: all four domains now appear by name — `libcucascade` showed as `domain0x1` before the ordering fix. - Builds clean after rebase onto `dev` (RAPIDS nightly 26.12): `sirius_extension`, `sirius_unittest`, `parquet_benchmark`, standalone `simpatico_codegen`. - `pre-commit run -a` green. ## Checklist - [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist - [ ] Cover changes with new or existing tests — NVTX domain membership and Quent capture timing are observable only in a profiler, so both are verified via the nsys/Quent runs above; the two `telemetry_context::create` call sites in tests are updated for the new signature - [x] Document configuration changes in code and summarize in the description above - [ ] Update human and agent documentation — `quent-telemetry.md` doesn't enumerate domains, so nothing went stale ## References Refs [cuCascade#196](NVIDIA/cuCascade#196), [quent#736](rapidsai/quent#736) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…irius-db#1809) Some clean up: - silence some build warnings - remove some obsolete rapids comparability branches
…s-db#1812) I was going through the simpatico README and found that (1) some of the commands were outdated and (2) I hit an RMM version bump-related error (`rmm::mr::set_current_device_resource_ref()` has been removed upstream in RMM). This PR fixes both these issues. Here's a concise summary of the build issue and proposed fix from Codex that I've edited myself > ### Build issue > The standalone Simpatico CLI and two tests failed against the upgraded RAPIDS/RMM stack because they still used the removed `rmm::mr::set_current_device_resource_ref()` API. The existing guards also saved the previous resource as a non-owning `device_async_resource_ref`, which can dangle after the current resource is replaced. > > ### Fix > Switch resource installation to `cudf::set_current_device_resource()`, retain its returned owning `cuda::mr::any_resource`, and move that resource back when restoring the previous allocator. This follows the ownership handling introduced in sirius-db#877, the non-deprecated cuDF API migration in sirius-db#1098, and the existing RAII pattern from sirius-db#1205. > The change here update the Simpatico CLI, operator-sweep test, and compression round-trip test. The build and CLI round-trip verification now pass. cc @felipeblazing --------- Signed-off-by: James Bourbeau <jbourbeau@nvidia.com>
## Description CTrack was fetched and linked unconditionally, which made a profiling-only dependency part of every build. This adds `BUILD_WITH_CTRACK`, defaulting to `OFF`, and gates the CTrack fetch, link dependency, and compile definition behind that option so normal builds do not require CTrack. ## Validation - `pixi run pre-commit run -a` passes at this commit. - The final stacked tree passes the complete release build. ## Checklist - [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements. - [x] Preserve the existing CTrack path when `BUILD_WITH_CTRACK=ON`. - [x] Document the new build option next to its CMake definition. ## Stack Layer 2 of 6. Based on `stacked/frugal-native-01-rapids`; followed by sirius-db#1776. --------- Co-authored-by: Matthijs Brobbel <m1brobbel@gmail.com>
## Description Replace the two watchdog runners' `fork()` + environment mutation + `exec()` sequences with `posix_spawn()` using an environment prepared entirely in the parent process. Add a reusable child-process environment helper that snapshots the inherited environment, applies child-only overrides and removals, and leaves the parent unchanged. This generalizes the existing envp-construction pattern in `test_operator_sweep.cpp` and `test_fused_operator_sweep.cpp` so both watchdogs share it. The spawn environments preserve the original pre-exec behavior, including inheriting `SIRIUS_DISABLE`. After re-exec, the hidden runners reapply `SIRIUS_CONFIG_FILE` and clear `SIRIUS_DISABLE` because shared test-environment startup resets those process-wide values before the selected test runs. A tree-wide process-launch audit found no other post-fork environment mutation. Other fork/exec sites either prepare `envp` in the parent or perform only descriptor and exec operations in the child, so they are intentionally outside this change. ## Validation Validated on Ubuntu 24.04 with an NVIDIA GeForce RTX 5060 Ti (compute capability 12.0), CUDA 13.2, and `io_uring` enabled: - `CUDAARCHS=120 CMAKE_BUILD_PARALLEL_LEVEL=8 pixi run make` - `pixi run pre-commit run -a` - Child-process environment regression: 7 assertions in 1 test case - Query-lifecycle watchdog: 78 assertions in 11 test cases - Hive-partition CUDA watchdog: 28 assertions in 1 test case - `pixi run make s3-test`: 31,433 assertions in 92 test cases using testcontainers-managed MinIO over HTTP and TLS - Native macOS C++17 spawn/environment smoke test passed ## Checklist - [x] Read CONTRIBUTING.md - [x] Cover changes with new or existing tests - [x] Document configuration changes in code and summarize in the description above (N/A: no configuration changes) - [x] Update human and agent documentation (N/A: internal test infrastructure only) ## References Closes sirius-db#1440 --------- Co-authored-by: Aaron Wu <Aaron.Wu@dell.com>
Reduce dependency on DuckDB-owned standard-library-like types. - Replace DuckDB vectors and smart pointers for Sirius pipeline/meta-pipeline ownership with standard-library equivalents. - Propagate the standard types through pipeline conversion, query indexing, task creation, telemetry, operator ports, and focused test helpers. - Remove direct DuckDB common-pointer/container includes where the pipeline headers no longer require them.
## Description There used to be separate compilation modes in Simpatico for decode vs encode, 1 of them using c++17 and the other c++20. There isn't really a reason to do so, and it turns out c++20 is also faster. ## Checklist - [ ] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist - [ ] Cover changes with new or existing tests - [ ] Document configuration changes in code and summarize in the description above - [ ] Update human and agent documentation (README.md, docs/, skills, CLAUDE.md) ## References Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
## Description Query lifecycle notifications were coupled to their producers, making it difficult for readahead and other services to observe execution without adding direct dependencies. This introduces a query event publisher and hook-driven subscribers, stamps publications with process-wide event IDs and system-clock timestamps, and adds the thread utilities used to dispatch subscriber work safely. Tests cover event publication, subscriber behavior, ordering metadata, and the supporting thread utilities. ## Validation - `pixi run pre-commit run -a` passes at this commit. - The final stacked tree passes the complete release build. ## Checklist - [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements. - [x] Cover the publisher, subscribers, and thread utilities with tests. - [x] No user-facing configuration change. ## Stack Layer 3 of 6. Based on `stacked/frugal-native-02-ctrack`; followed by sirius-db#1774.
…ade/offload copies (sirius-db#1552) # spill: fix the partition-spill cliff and monolithic downgrade copies (T8) **Branch:** `pr/spill-cliff-chunked-offload` → base: `dev` (`3a0ef61f`). No stacking dependency. ## Problem The q9-class RF1 spill cliff: +0.1 % delta rows tipped two extra ~4.5 GB partition round-trips through host (D2H 57→75 GB / H2D 75→93 GB, post-RF1 +117 ms), delivered as single 18–28 GB blocking copy submissions on an un-instrumented thread. Two independent causes: 1. The builtin cucascade GPU→HOST fast converter submits an entire batch's D2H copies as one monolithic batched call plus a blocking sync. 2. Monitor-issued downgrades force-flush the whole trigger→stop band (≥ 12 GB per marginal crossing on the SF1000 GB300 config) and spill whole partitions in policy order, so a marginal working-set increase becomes a multi-GB spill. ## Change 1. **Chunked, pipelined GPU→HOST spill copies** (`src/data/spill_chunked_converters.cpp`, `chunked_spill_copy.hpp`): `SiriusContext::initialize` replaces the builtin converter with a Sirius-side transcription that flushes D2H submissions every `~copy_chunk_bytes` (default 1 GiB; yaml `executor.downgrade.copy_chunk_bytes`, `0` keeps the builtin) while the column tree is still being walked — chunk k's DMA overlaps chunk k+1's prep. Produces a byte-identical `host_data_representation` (same self-describing column_metadata layout); the HOST→GPU restore direction deliberately stays on the unchanged builtin converter. Adds NVTX ranges to the downgrade worker's conversions (no timing instrumentation is compiled into the production path; measurement is via the profiler). 2. **Overflow-proportional spill sizing** (`downgrade_executor.cpp`, new `spill_policy.hpp`): - Monitor requests (yaml `executor.downgrade.overflow_proportional_spill`, default true) stop as soon as live pressure drops back below the *trigger* threshold instead of flushing down to the stop threshold. - Byte-targeted requests carry `target_bytes`: dispatch stops once planned (freed + in-flight) bytes cover the target, bounding overshoot to < 1 batch at any pool width, and the final per-repo pick is best-fit (smallest candidate covering the remaining deficit) instead of the next whole partition in policy order. - The live predicate is evaluated before every dispatch, so an already-satisfied request spills nothing instead of at least one batch. Reservation-driven `request_downgrade` keeps its dispatch-until-satisfied envelope (livelock-sensitive at 251/256 GB); it only gains the earlier predicate checks. This decouples spilled bytes from both the trigger→stop band and whole-partition granularity **without** retuning `hash_partition_bytes` (a measured trade, not a win). ## Measured SF1000, GB300: **q9 post-RF1 spill cliff eliminated (+117 ms → −71 ms)**; spilled bytes become proportional to the actual overflow instead of the trigger→stop band. ## Testing - New unit tests: `test_chunked_spill_copy.cpp` (chunked converter produces byte-identical host representations across chunk-size sweeps), `test_spill_policy.cpp` (best-fit picks), extended `test_downgrade_executor.cpp` (predicate-stops, target_bytes bounding). - Full `sirius_unittest` suite green on GB300: 2842 test cases, 32,778,745 assertions. - Full pre-commit hook set passes on the changed files. ## Related follow-up (not in this PR) The stack lineage carries `dcc4c50a` "fix(downgrade): wake the scheduler matcher after returning extracted tasks" — a pre-existing whole-process deadlock (task returned via RAII emits no `task_available`; deterministic at SF1000 host-pinned q18, 7/7 hangs). It applies to `dev` independently of this PR and should be its own small PR. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Bobbi Winema Yogatama <34973829+bwyogatama@users.noreply.github.com>
…g, clause-2.8 evolution, scratch rollback, nsys mapping) (sirius-db#1554) # bench: official TPC-H power/throughput/QphH harness (RF1/RF2, per-query telemetry + nsys) Tooling-only PR — no engine code. Branch: `pr/tpch-power-throughput-harness`, rebased onto current `dev` (2026-09-16). ## Description Extends `test/tpch_performance/tpch_power_throughput.py` and the `bench/sf1000-repro` kit into a harness for **official TPC-H power/throughput scoring** on a file-backed `.duckdb` dataset, with per-query telemetry and per-query nsys profiling for analysis runs. This harness produced the SF1000 QphH campaign on GB300 that took **QphH@Size from 5.47M to 7.52M** (reference run in the commit message: Power@Size 7,165,616 / Throughput@Size 4,951,261 / QphH 5,956,411 on the PR sirius-db#1409 stack, GPU-vs-CPU validation PASS after both refresh functions). ### Scoring (spec-conformant) - `Power@Size = 3600·SF / geomean(22 stream-0 query times + T_RF1 + T_RF2)` — the refresh functions are timed as spec elements of the power geomean, not warmup. Query times below `slowest/1000` are floored per **clause 5.4.1.4** (the 1000:1 spread cap). - `Throughput@Size = N·22·3600 / measurement_interval · SF` — the interval is wall-clock from first stream start to last stream end (**clause 5.3.6.1**), with N concurrent query streams (spec permutations) plus one refresh stream running N RF1/RF2 pairs. - `QphH@Size = sqrt(Power · Throughput)`. - `--vary-predicates` runs each stream's own qgen substitution parameters, as an official run requires; the default fixed parameters keep validation meaningful. - `--update-set-offset` implements **clause 2.8 database evolution**: one evolving database is reused across runs (`--scratch-db <shared> --keep-scratch-db`) — the power run consumes update set N+1, throughput N+2..N+streams+1 — instead of restoring a 440 GB pristine copy per benchmark run. Validation requires offset 0 (`--update-set-offset` rejects `--validation`); the scratch size check accepts the evolved size and reports drift instead of failing. ### Shared-GPU probe gate (`run-power.sh`) The config reserves ~0.95 of the device at `LOAD`. On a shared box that fails not only while another workload runs but **for minutes after it exits** — the driver reclaims a dead process's async-pool backing lazily, during which `nvidia-smi` already reports the memory free but large allocations still OOM. Gating on `nvidia-smi` therefore passes and the run still dies; the only reliable gate is the failing operation itself. `run-power.sh` retries a throwaway extension-LOAD probe (`PROBE_TRIES`, default 20 × 60 s) before committing to the run, on the original (non-telemetry) config so a Quent capture never sees a second same-named engine. ### Rollback-scratch: WAL-discard restore (`--rollback-scratch`, `ROLLBACK=1`) Restores the scratch DB after a run **without re-copying** (a 15-minute, 440 GB operation at SF1000). Recipe — three thresholds keep every refresh mutation confined to the WAL: `SET checkpoint_threshold='1TB'`, `SET wal_autocheckpoint='1TB'`, and `SET auto_checkpoint_skip_wal_threshold=2^40` (so bulk writes do not checkpoint-on-commit). The run then ends via `os._exit` because DuckDB force-checkpoints the WAL into the base file on any clean shutdown and offers no opt-out; deleting `<scratch>.wal` afterwards (run-power.sh's `ROLLBACK=1` does it) restores content-pristine state, so every run reuses one copy at offset 0. Qualified at small scale: 3 mutate/discard cycles returned exact pristine content each time. **Caveat (orphan drift):** the base file keeps small orphaned-block growth from DuckDB's optimistic bulk writes — never read, never reclaimed — so the file size drifts up across cycles; the size check reports the drift, and the copy should be refreshed when it grows large. Incompatible with `NSYS=1`. ### Per-query nsys (`--nsys-per-query`, `NSYS=1`) Each sequential power-pass query is bracketed in its own cudaProfilerApi capture range (`--capture-range=cudaProfilerApi --capture-range-end=repeat::sync`); the throughput phase gets one whole-interval range (cudaProfilerStart is process-global — concurrent streams must not use it). Two gotchas are baked in: - **`repeat::sync`, and never trust report indices**: nsys silently merges adjacent capture ranges when profiler_stop/start pairs arrive faster than it finalizes a range, and can drop the last range at exit — the surviving `range.<N>.nsys-rep` files are numbered compactly, so index-based assignment shifts every pointer after the first merge (observed: 72 files for 89 planned ranges). `prep_analysis_bundles.py` therefore joins reports to manifest entries by **absolute-timestamp overlap** (report capture window from its sqlite export, manifest window from the runner-recorded `start/stop_epoch_ns`, falling back to the quent query's Init..Exit window), and writes a full accounting to `nsys_range_map.json` — mapped / ambiguous (merged) / dropped / no_window / conflict; orphan reports and partial overlaps flagged, nothing silently reassigned. A 0-byte sqlite export is treated as failed, not cached. - **env-clear for the nsys frontend**: the frontend's bundled `libssl.so.3` predates the pixi libcurl's `OPENSSL_3.2.0` requirement and dies at startup, so nsys is launched with `LD_PRELOAD`/`LD_LIBRARY_PATH` cleared and both restored only for the profiled python. nsys runs are analysis runs; their metrics are never quoted as scores. ### Validation worker design Validation cannot share a process with the pinned run: Sirius's host pool is a growing pool allocator (unpinning returns blocks to the pool, not the OS) and DuckDB sizes its default `memory_limit` from total system RAM with no knowledge of what Sirius holds — at SF1000 a CPU-side q9 allocates into a machine already ~320 GB spoken for and gets OOM-killed. The GPU rows are pickled per phase, and a **fresh child process** that never loads the extension replays RF1/RF2 on its own copy of the pristine input to reproduce the post-refresh states, then diffs row-by-row (sorted multisets, absolute tolerance). Related: `--duckdb-memory-limit` (default 32 GB) caps the benchmark connection itself — the unlimited default OOM-killed the pinned process during pin materialization at SF1000. ### Telemetry and analysis prep - Every query is labeled (`CALL sirius_set_query_label`, zero-padded `q01`..`q22`) and each phase bucketed into its own Quent query group via `CALL sirius_set_session_label` (`warmup`, `power_clean`, `power`, `power_postrf2`, `tput_s1..sN`); both calls sit outside the timed window. `QUENT=1` derives a telemetry-enabled config per run. - `prep_analysis_bundles.py <run_dir> <nsys_dir> <quent_dir>` builds per-query analysis bundle JSONs (nsys report pointers with `nsys_status`/`nsys_shared_with`, timings, quent UUIDs) and pre-exports every `.nsys-rep` to sqlite in parallel. - `--pin-layout <json>`: mixed-tier pin layouts with validated column coverage (a schema-qualified name creates a second entry over the same table for per-column tiering); SF1000 reference layout in `bench/sf1000-repro/pin-layout-sf1000.json`. - `--warmup-pass` burns one discarded pass so JIT/first-touch costs land nowhere; crash-safe cleanup keeps results when unpin fails; the perf-stack env (`SIRIUS_EXP_*`, `SIRIUS_PRE_SQL`, `LD_PRELOAD`) is recorded into `run_info.txt`/`metrics.json` for provenance. - `generate_tpch_refresh.sh`: dbgen `upath[128]` overflow fix (stage via `/tmp`). ## Engine dependencies (all in `dev` now) Everything this harness leans on has landed, and the rebase commit adapts the harness to it: - **`sirius_set_session_label`** exists in `dev`; the runner still catches `CatalogException` and degrades to unlabeled groups on an older engine. - **Split pin entries** (the `main.orders` host-tier entry in the SF1000 layout) use the column-aware `find_pinned_entry_for_duckdb_table(requested_ids)` lookup that is in `dev`. A scored run must still have zero `Transparent execution fallback` lines in its `log_dir`. - **`SIRIUS_EXP_*` gates** (sirius-db#1524): `run-power.sh` exports the same names and defaults as `run.sh` (`SIRIUS_EXP_FUSED_SCAN_FILTER`, `SIRIUS_EXP_LATE_MAT`, `SIRIUS_EXP_LATE_MAT_PIN_UNIQUE_COLS`). Late-mat is inert on duckdb pins; fused scan-filter engages on the compressed GPU pins. - **Config** (sirius-db#1499): the two diagnostic YAMLs no longer set `thread_name_prefix`, which the loader now rejects. - **Patched libcudf** is optional (`CUDF_SO`), matching `run.sh` after sirius-db#1524; unset uses the pixi libcudf. - `SIRIUS_POISON_FREES` / `SIRIUS_QUARANTINE_FREES` (used by `verify-poison-concurrent.sh`) are the one remaining separate piece; without them that gate degrades to a concurrent stress run with coredump capture. ## Reconciliation with sirius-db#1470 PR sirius-db#1470 (fadvise cache eviction, benchmark totals) landed in `test/tpch_performance/performance_test.py`; this PR does not touch that file and keeps dev's shape. The runner's own `_evict_page_cache` uses the same mechanism (`posix_fadvise(DONTNEED)`, no sudo) but targets the `.duckdb` scratch/input files with an fsync-first dirty-page path — it complements rather than duplicates `drop_os_cache`. ## Staged refresh (sirius-db#1551) Staged refresh landed in `dev` while this PR was open and is on by default in the runner; the clause-2.8 offset and the rollback path both pass it through (`rf1/rf2_statements(dir, set, staged)`), and `run-power.sh` forwards `--no-staged-refresh` for the legacy path. Staging tables are created after the rollback thresholds are armed, so they live in the WAL and are discarded with it. ## Rollback on every exit path The rebase commit also closes a gap in `--rollback-scratch`: an exception mid-run used to reach interpreter teardown, where the kept-alive connection's destructor checkpoints the WAL into the base file and consumes the scratch — on exactly the crashed run whose scratch needs restoring. `os._exit` now fires on every exit path once rollback is armed, and `run-power.sh` reaches its WAL-discard step even when the runner fails (`|| RC=$?` instead of tripping `set -e`). ## Checklist - [x] Read CONTRIBUTING.md - [x] Cover changes with new or existing tests — `py_compile`/`bash -n`/`pre-commit` clean; exercised end-to-end at SF1000 (see measured context above) and at small scale for the rollback qualification. No unit tests: benchmark tooling, not engine code. - [x] Document configuration changes — knob table in `bench/sf1000-repro/README.md`, flag docs in `test/tpch_performance/CLAUDE.md`. - [x] Update human and agent documentation (README.md, CLAUDE.md) ## References Refs sirius-db#1470 (cache-eviction overlap, reconciled above). 🤖 Generated with [Claude Code](https://claude.com/claude-code) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…irius-db#1816) As mentioned in sirius-db#1806 (comment), these files are obsolete.
## Description Components that finish work on CUDA streams need to observe event completion without blocking an executor thread. This adds `cuda_event_completion_poll`, which polls CUDA events and dispatches their completion callbacks through the execution infrastructure. The accompanying tests cover registration, polling, completion dispatch, shutdown, and error paths. ## Validation - `pixi run pre-commit run -a` passes at this commit. - The final stacked tree passes the complete release build. ## Checklist - [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements. - [x] Cover normal completion and lifecycle edge cases with tests. - [x] No user-facing configuration change. ## Stack Layer 4 of 6. Based on `stacked/frugal-native-03-query-events`; followed by sirius-db#1773.
…#1773) ## Description The I/O code used prefetch-specific names for objects that now represent cache residency and completion more generally. This normalizes the I/O namespace and source layout, moves REST and S3 helpers under `io/rest`, introduces the shared invocable execution utility, and renames `prefetching_handle` to `cache_handle` throughout scan, cache, and engine code. The rename prepares the interfaces consumed by the dynamic I/O and frugal-caching layer while keeping behavior unchanged. ## Validation - `pixi run pre-commit run -a` passes at both commits in this PR. - The final stacked tree passes the complete release build. ## Checklist - [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements. - [x] Update affected tests and documentation with the normalized names. - [x] No user-facing configuration change. ## Stack Layer 5 of 6. Based on `stacked/frugal-native-04-cuda-events`; followed by sirius-db#1771.
…#1593) This PR simply updates sirius to be able to compile against updates in NVIDIA/cuCascade#192, so we can incrementally move to a more bolted down data management strategy. This is intentionally narrowly scoped to ease the review process, and the subsequent independent pieces will be proposed separately. Part 1 of a few to move towards solving sirius-db#1063 from the root of it.
…sirius-db#1828) Bumps [clap](https://github.com/clap-rs/clap) from 4.6.6 to 4.6.7. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/clap-rs/clap/releases">clap's releases</a>.</em></p> <blockquote> <h2>v4.6.7</h2> <h2>[4.6.7] - 2026-09-14</h2> <h3>Features</h3> <ul> <li><em>(derive)</em> Add <code>#[command(defer = <bool>)]</code> attribute to opt-in to lazy initialisation of subcommands</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/clap-rs/clap/blob/main/CHANGELOG.md">clap's changelog</a>.</em></p> <blockquote> <h2>[4.6.7] - 2026-09-14</h2> <h3>Features</h3> <ul> <li><em>(derive)</em> Add <code>#[command(defer = <bool>)]</code> attribute to opt-in to lazy initialisation of subcommands</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/clap-rs/clap/commit/d3e59a9ab214910b9dad02921b7ef42c6400de9b"><code>d3e59a9</code></a> chore: Release</li> <li><a href="https://github.com/clap-rs/clap/commit/d997f878c484e02b2935d32a7b67afe990e91227"><code>d997f87</code></a> docs: Update changelog</li> <li><a href="https://github.com/clap-rs/clap/commit/fb6058cc38cf6f4b07b8db89c850fc80e059e222"><code>fb6058c</code></a> Merge pull request <a href="https://redirect.github.com/clap-rs/clap/issues/6409">#6409</a> from heaths/pwsh-support</li> <li><a href="https://github.com/clap-rs/clap/commit/2310870f7a9fe4c39f5bb5620261a06ac9aaa0b1"><code>2310870</code></a> test(complete): Add tests for completer_for_path</li> <li><a href="https://github.com/clap-rs/clap/commit/5967c17393a09f90112cdc5b2c20a1383f0bd8e4"><code>5967c17</code></a> refactor(complete): Move shell detection to Shells</li> <li><a href="https://github.com/clap-rs/clap/commit/594602bb29f374e49da5029df63e5c740c723ad0"><code>594602b</code></a> fix(complete): Detect pwsh for PowerShell</li> <li><a href="https://github.com/clap-rs/clap/commit/3a4f2d031b5110aaaac40f1cb1d6c0b8ff619df8"><code>3a4f2d0</code></a> Merge pull request <a href="https://redirect.github.com/clap-rs/clap/issues/6427">#6427</a> from clap-rs/renovate/shlex-2.x</li> <li><a href="https://github.com/clap-rs/clap/commit/67ebaed4ac95a83d5d7cc22d64ee0fc162ae1db3"><code>67ebaed</code></a> Merge pull request <a href="https://redirect.github.com/clap-rs/clap/issues/6426">#6426</a> from clap-rs/renovate/actions-checkout-7.x</li> <li><a href="https://github.com/clap-rs/clap/commit/c968b136d0e6f560c08b06a844a20f2798a00096"><code>c968b13</code></a> chore(deps): Update Rust crate shlex to v2</li> <li><a href="https://github.com/clap-rs/clap/commit/8f247cbf58c582a755ee4dbe62a2390d93b9f8eb"><code>8f247cb</code></a> chore(deps): Update actions/checkout action to v7</li> <li>Additional commits viewable in <a href="https://github.com/clap-rs/clap/compare/clap_complete-v4.6.6...clap_complete-v4.6.7">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [clap](https://github.com/clap-rs/clap) from 4.6.6 to 4.6.7. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/clap-rs/clap/releases">clap's releases</a>.</em></p> <blockquote> <h2>v4.6.7</h2> <h2>[4.6.7] - 2026-09-14</h2> <h3>Features</h3> <ul> <li><em>(derive)</em> Add <code>#[command(defer = <bool>)]</code> attribute to opt-in to lazy initialisation of subcommands</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/clap-rs/clap/blob/main/CHANGELOG.md">clap's changelog</a>.</em></p> <blockquote> <h2>[4.6.7] - 2026-09-14</h2> <h3>Features</h3> <ul> <li><em>(derive)</em> Add <code>#[command(defer = <bool>)]</code> attribute to opt-in to lazy initialisation of subcommands</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/clap-rs/clap/commit/d3e59a9ab214910b9dad02921b7ef42c6400de9b"><code>d3e59a9</code></a> chore: Release</li> <li><a href="https://github.com/clap-rs/clap/commit/d997f878c484e02b2935d32a7b67afe990e91227"><code>d997f87</code></a> docs: Update changelog</li> <li><a href="https://github.com/clap-rs/clap/commit/fb6058cc38cf6f4b07b8db89c850fc80e059e222"><code>fb6058c</code></a> Merge pull request <a href="https://redirect.github.com/clap-rs/clap/issues/6409">#6409</a> from heaths/pwsh-support</li> <li><a href="https://github.com/clap-rs/clap/commit/2310870f7a9fe4c39f5bb5620261a06ac9aaa0b1"><code>2310870</code></a> test(complete): Add tests for completer_for_path</li> <li><a href="https://github.com/clap-rs/clap/commit/5967c17393a09f90112cdc5b2c20a1383f0bd8e4"><code>5967c17</code></a> refactor(complete): Move shell detection to Shells</li> <li><a href="https://github.com/clap-rs/clap/commit/594602bb29f374e49da5029df63e5c740c723ad0"><code>594602b</code></a> fix(complete): Detect pwsh for PowerShell</li> <li><a href="https://github.com/clap-rs/clap/commit/3a4f2d031b5110aaaac40f1cb1d6c0b8ff619df8"><code>3a4f2d0</code></a> Merge pull request <a href="https://redirect.github.com/clap-rs/clap/issues/6427">#6427</a> from clap-rs/renovate/shlex-2.x</li> <li><a href="https://github.com/clap-rs/clap/commit/67ebaed4ac95a83d5d7cc22d64ee0fc162ae1db3"><code>67ebaed</code></a> Merge pull request <a href="https://redirect.github.com/clap-rs/clap/issues/6426">#6426</a> from clap-rs/renovate/actions-checkout-7.x</li> <li><a href="https://github.com/clap-rs/clap/commit/c968b136d0e6f560c08b06a844a20f2798a00096"><code>c968b13</code></a> chore(deps): Update Rust crate shlex to v2</li> <li><a href="https://github.com/clap-rs/clap/commit/8f247cbf58c582a755ee4dbe62a2390d93b9f8eb"><code>8f247cb</code></a> chore(deps): Update actions/checkout action to v7</li> <li>Additional commits viewable in <a href="https://github.com/clap-rs/clap/compare/clap_complete-v4.6.6...clap_complete-v4.6.7">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
I was going through the Python API section of the README and noticed it
fails to build due to lack of a recent git tag
<details>
<summary>Full traceback:</summary>
```
(base) jbourbeau@rl-b5652-03:~/sirius-python-build$ pixi run -e duckdb-python build-duckdb-python
✨ Pixi task (build-duckdb-python in duckdb-python): CMAKE_ARGS="-DDUCKDB_SOURCE_PATH=$PIXI_PROJECT_ROOT/duckdb" pip install --no-build-isolation ./duckdb-python
Processing ./duckdb-python
Preparing metadata (pyproject.toml) ... error
error: subprocess-exited-with-error
× Preparing metadata (pyproject.toml) did not run successfully.
│ exit code: 1
╰─> [77 lines of output]
[version_scheme] version object: <ScmVersion 0.0.1.dev1 dist=1 node=gb236c8194ed14c7a7c685e0534dde501cc855b3a dirty=False branch=HEAD>
[version_scheme] version.tag: 0.0.1.dev1
[version_scheme] version.distance: 1
[version_scheme] version.dirty: False
Traceback (most recent call last):
File "/home/jbourbeau/sirius-python-build/duckdb-python/duckdb_packaging/setuptools_scm_version.py", line 56, in version_scheme
return _bump_dev_version(str(version.tag), distance)
File "/home/jbourbeau/sirius-python-build/duckdb-python/duckdb_packaging/setuptools_scm_version.py", line 73, in _bump_dev_version
major, minor, patch, post, rc = parse_version(base_version)
~~~~~~~~~~~~~^^^^^^^^^^^^^^
File "/home/jbourbeau/sirius-python-build/duckdb-python/duckdb_packaging/_versioning.py", line 33, in parse_version
raise ValueError(msg)
ValueError: Invalid version format: 0.0.1.dev1 (expected X.Y.Z, X.Y.Z.rcM or X.Y.Z.postN)
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/pip/_vendor/pyproject_hooks/_in_process/_in_process.py", line 389, in <module>
main()
~~~~^^
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/pip/_vendor/pyproject_hooks/_in_process/_in_process.py", line 373, in main
json_out["return_val"] = hook(**hook_input["kwargs"])
~~~~^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/pip/_vendor/pyproject_hooks/_in_process/_in_process.py", line 175, in prepare_metadata_for_build_wheel
return hook(metadata_directory, config_settings)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/build/__init__.py", line 96, in prepare_metadata_for_build_wheel
return _build_wheel_impl(
~~~~~~~~~~~~~~~~~^
None, config_settings, metadata_directory, editable=False
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
).wheel_filename # actually returns the dist-info directory
^
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/build/wheel.py", line 177, in _build_wheel_impl
return _build_wheel_impl_impl(
wheel_directory,
...<5 lines>...
pyproject=pyproject,
)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/build/wheel.py", line 234, in _build_wheel_impl_impl
metadata = get_standard_metadata(pyproject, settings)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/build/metadata.py", line 44, in get_standard_metadata
new_pyproject_dict["project"] = process_dynamic_metadata(
~~~~~~~~~~~~~~~~~~~~~~~~^
new_pyproject_dict["project"], settings.metadata
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/builder/_load_provider.py", line 150, in process_dynamic_metadata
return dict(settings)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/builder/_load_provider.py", line 114, in __getitem__
self.project[key] = provider.dynamic_metadata( # type: ignore[call-arg]
~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
key, self.settings[key]
^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/metadata/setuptools_scm.py", line 30, in dynamic_metadata
version = _get_version(config, force_write_version_files=True)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/setuptools_scm/_get_version.py", line 29, in _get_version
return _get_version_core(
config, force_write_version_files=force_write_version_files
)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/vcs_versioning/_get_version_impl.py", line 171, in _get_version
return _finalize(scm_version, config, force_write=force_write_version_files)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/vcs_versioning/_get_version_impl.py", line 53, in _finalize
version_string = _format_version(applied)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/vcs_versioning/_version_schemes/__init__.py", line 113, in format_version
main_version = _entrypoints._call_version_scheme(
version_for_scheme,
"setuptools_scm.version_scheme",
version.config.version_scheme,
)
File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/vcs_versioning/_entrypoints.py", line 87, in _call_version_scheme
result = scheme(version)
File "/home/jbourbeau/sirius-python-build/duckdb-python/duckdb_packaging/setuptools_scm_version.py", line 59, in version_scheme
raise RuntimeError(msg) from e
RuntimeError: Failed to bump version: Invalid version format: 0.0.1.dev1 (expected X.Y.Z, X.Y.Z.rcM or X.Y.Z.postN)
[end of output]
note: This error originates from a subprocess, and is likely not a problem with pip.
error: metadata-generation-failed
× Encountered error while generating package metadata.
╰─> from file:///home/jbourbeau/sirius-python-build/duckdb-python
```
</details>
This PR removed `--depth=1` in the submodule init command to ensure we
pull the full git history (including tags). I also added a small CI
build step where we ensure the Python API builds successfully (similar
conditions as the rust binding checks).
Signed-off-by: James Bourbeau <jbourbeau@nvidia.com>
Migrate to `cuda::stream_ref`. This fixes our nightly build and also gets rids of the deprecation warnings.
## Description At scale factor 3000 on a smaller machine, the new dense join operator (sirius-db#1606) goes off the rails and tries to reserve 800GB of device memory (10x the size of the device). This PR changes the estimation to hopefully track actual usage a bit closer (please have a look if you think there are situations where the estimate is now too low!). And it adds a fallback path for when the new operator will not fit. Adds config parameter dense_count_join_memory_fraction to set what part of the memory to use for the dense count join. Default 10%. The other parameter, setting an explicit byte count for the threshold, still exists but is now set to 0 (auto) to use the new fraction. ## Checklist - [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist - [x] Cover changes with new or existing tests - [x] Document configuration changes in code and summarize in the description above - [x] Update human and agent documentation (README.md, docs/, skills, CLAUDE.md) ## References --------- Co-authored-by: Joost Hoozemans <1442581+joosthooz@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
We can unify our instructions now that claude also reads them.
This allows us to drop the `protoc` dependency requirement.
## Description Add GPU scan support for the following Parquet virtual columns: - `filename` - `file_index` - `file_row_number` This PR: - preserves file identity and file-local row offsets in Parquet split metadata; - reads selected row groups as provenance-preserving runs, synthesizes the requested virtual columns on GPU, and concatenates the results; - reuses the Parquet batch-layout calculation for both `file_row_number` and Iceberg positional-delete processing; - evaluates predicates involving virtual columns after virtual-column materialization; - disables row-level cuDF predicate pushdown when virtual columns are requested, while retaining physical row-group statistics pruning; - supports virtual-only scans, projection/reordering, duplicate outputs, Hive partition columns, legacy virtual-column options, and Iceberg scans; - bypasses pinned Parquet cache entries for virtual-column queries because existing pins do not store per-row provenance. There are no new configuration options. ### file_index compatibility Sirius keeps `file_index` as the zero-based index in the original bound file list. For a bound list `[A, B]`, B retains index 1 when a runtime join filter removes A. DuckDB can renumber B to 0 when `file_index` is projected and runtime filtering targets either a Hive partition column or the legacy filename column exposed by `filename=true`. The standard virtual `filename` column alone does not enable this pruning. The upstream inconsistency is tracked in duckdb/duckdb#26044 and is also reproducible in the standard DuckDB CLI v2.0.0-dev3 (`ed599a4425`). The runtime-pruning compatibility test covers both triggers. If DuckDB adopts stable numbering, the test should require CPU/GPU equality and remove its known-difference expectation. ### Follow-up work - Use cuDF 26.08.01's existing `enable_prepend_source_index_column()` and `enable_prepend_row_index_column()` APIs to read ordinary Parquet splits in one invocation while retaining physical reader filtering and dynamic-filter AST merging. Map source indices to bound file identities, adapt the column layout, convert row indices to BIGINT, and update memory accounting and provenance tests. - Translate predicates on `filename` and `file_index` into source selection. - Adapt Iceberg positional-delete matching to mapped file identities and original row indices before enabling reader filtering for that path. - Add explicit `pin_table` support for Parquet virtual columns. - Improve physical carrier selection and normalization, and determine when virtual-only scans can safely avoid physical reads. - Include the filename scalar's device buffer (`path.size()` bytes) in construction memory accounting. ### Validation Previously reported local validation (not a final-head CI result): ```text [scan][parquet][assembly] 6 test cases, 36 assertions passed [integration][gpu_execution][scan][virtual_columns][acceptance] 8 test cases, 7195 assertions passed [integration][gpu_execution][hive_partition] 7 test cases, 366 assertions passed ``` The full C++ test suite was also run. Two unrelated REST subprocess tests initially hit local GPU OOM and both passed on retry. The 14 previously hidden integration cases are now visible to the default test run. Final-head CI must pass and exercise them, the runtime-pruning compatibility case, and the filename sizing cases. ## Checklist - [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist - [x] Cover changes with new or existing tests - [x] Document configuration changes in code and summarize in the description above — no configuration changes - [x] Update human and agent documentation — no new configuration or agent-facing interface ## References Refs sirius-db#492
…#1874) ## Description Fix the arm64 nightly build failure introduced by sirius-db#1846. The new Parquet virtual-column helpers declared `rmm::cuda_stream_view` without including its declaration, causing compilation to fail in `scan_plan.hpp` and `scan_plan.cpp`. Use `cuda::stream_ref` in the helpers and their callers, matching the surrounding scan API. Update the scan output assembly test helper to use the same type. Verification: - Compiled the affected scan objects and scan output assembly test source on x86. - Commit hooks passed, including clang-format. - The arm64 nightly build has not been run locally; CI will verify that target. ## Checklist - [x] Read CONTRIBUTING.md and ensure PR meets the reviewability checklist - [x] Cover changes with existing scan tests - [x] Document configuration changes in code and summarize them above — no configuration changes - [x] Update human and agent documentation — no documentation changes needed for this build fix ## References - Follow-up to sirius-db#1846
This adds the `sirius` Rust crate docs to our GitHub pages deploy (in addition to sirius-db#1860).
…-db#1877) ## Description Land the prep work sirius-db#1859's rename cutover depends on, so main-triggered CI and contributor-facing docs are correct before the freeze starts: CI workflows now trigger on main alongside dev, CONTRIBUTING.md gains a local-clone migration guide covering every remote layout in use (personal fork, Stacked-PR-maintainer origin, fallback), dry-run verified against real git behavior, and stray dev-branch references across docs and one source comment now point at main. **NOTE:** There will be a follow-up PR once the rename happens on `main` to remove any CI references for `dev`. These are needed with `main` to work across the rename ## Checklist - [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist - [ ] Cover changes with new or existing tests - [ ] Document configuration changes in code and summarize in the description above - [x] Update human and agent documentation (README.md, docs/, skills, CLAUDE.md) ## References Refs sirius-db#1859
…ius-db#1878) ## Description Follow-up to sirius-db#1877. The dev-to-main rename (sirius-db#1859) is complete and `dev` no longer exists as a branch, so the temporary dual `dev`/`main` CI triggers added as prep work are no longer needed. `check.yml`, `test.yml`, `distribution.yml`, and `docs.yml` now trigger on `main` only, and the two branch-name conditionals in `distribution.yml` and `docs.yml` drop their `dev` half. A repo-wide sweep found no other leftover `dev`-as-branch language that needs cleanup — the remaining mentions in `CONTRIBUTING.md`'s migration guide and `AGENTS.md`'s pointer sentence are intentional and still needed until every contributor's local clone/fork/stack has migrated. **NOTE:** This PR is non-blocking so it can be merged once it has been reviewed ## Checklist - [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist - [ ] Cover changes with new or existing tests - [ ] Document configuration changes in code and summarize in the description above - [x] Update human and agent documentation (README.md, docs/, skills, CLAUDE.md) ## References Closes sirius-db#1859
…og lowering/plan/execute timings (sirius-db#1841) ## Why Three gaps on the FFI path (`sirius_ffi.cpp`, the embedded engine used by `rust/crates/sirius-sys` and the `experimental/` backends), all found while profiling TPC-H SF100 through it: 1. **Substrait lowering re-parsed parquet footers on every bind.** The Substrait consumer builds the plan through the Relation API, which re-binds the whole subtree at every level, so a parquet read is bound once per operator above it. With the metadata cache off each bind re-parses the footer and re-derives the column statistics. On a 25 GB, ~5k-row-group SF100 `lineitem`, lowering took 1.3 s for Q6 and **10 s for Q21 — as long as or longer than the GPU execution (1.3 s / 4.6 s)**. 2. **No log output.** The embedded DuckDB never loads the Sirius extension, so nothing on this path read `SIRIUS_LOG_{BACKEND,DIR,LEVEL}` the way `SiriusContextExtensionCallback` does on the transparent path; the engine ran with the noop sink and the host process could not get a Sirius log. 3. **No per-phase timing.** The host only sees the total `execute_substrait` time, and the telemetry query window covers execution alone, so time spent in lowering/planning (like the 10 s above) was invisible. ## What - `bring_up()`: set `parquet_metadata_cache = true` on the embedded DuckDB **at the database level** (`DBConfig::SetOptionByName`). It has to be database-level: the consumer's binder does not see session-level `SET` variables. This DuckDB instance is private to the engine and the cache validates file mtimes, so there is no staleness hazard for the host. Lowering drops to 55–220 ms on the cases above. - `install_log_sink_from_env()`: at bring-up, if any of `SIRIUS_LOG_BACKEND` / `SIRIUS_LOG_DIR` / `SIRIUS_LOG_LEVEL` is set, copy them into `duckdb::Config::LOG_*` and call `install_configured_log_sink(nullptr)` (best-effort, an unknown backend is ignored rather than failing bring-up). With none of the variables set the sink is left untouched, i.e. behaviour is unchanged by default. - `execute_substrait()`: time the lowering / physical-planning / execution phases and emit one `SIRIUS_LOG_INFO` line per query. +52/−1, `sirius_ffi.hpp` unchanged, no API change. Applies to every FFI consumer (StarRocks backend, `sirius-sys`), all in the "faster and observable" direction. ## Verification - TPC-H SF1 22/22 through the FFI path unchanged vs the DuckDB baseline. - SF100 on an L40S (`g6e.8xlarge`): per-query engine time drops by the lowering delta (Q21 10 s → ~0.2 s lowering); the new log line is what exposed the split. - Existing `test/cpp/exec/test_sirius_ffi_fragment.cpp` covers bring-up/execute on this path; no new test since the change is configuration + logging. ## Relationship to other PRs Independent of the join fix (sirius-db#1840). sirius-db#1842's benchmark numbers were taken with this change in place. --------- Co-authored-by: morningman <moringman@apache.org> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ize inequality sides (sirius-db#1840) ## Why Two failures that only the Substrait/FFI path (`sirius_ffi.cpp` → `sirius_physical_plan_generator`) reaches. The transparent DuckDB-extension path plans the same SQL differently — DELIM_JOIN keeps the inequality in a filter, and the optimizer puts the `NOT IN` mark join on top — so it never hit them. Found by running TPC-H through the FFI path (the plans an Apache Doris FE produces, see sirius-db#1842) on a T4. ## What **`src/op/sirius_physical_hash_join.cpp` — MARK join polled before sizing (TPC-H Q16, process abort).** A MARK join whose probe pipeline is the source of an inner join's probe side gets polled by the task creator (`WAITING_FOR_INPUT_DATA` walk) before either of its input partitions negotiated `BUILD_PROBE`. `get_next_task_hint` then takes the STANDARD/MIXED branch and `refresh_cross_schedule` throws its MARK invariant from the task-creator thread, which terminates the process. A MARK join always ends up in `BUILD_PROBE`, so when polled in that window there is nothing to schedule yet: return `WAITING_FOR_INPUT_DATA` on the build producer instead, whose tasks feed the build partition and size the join. **`src/planner/sirius_plan_comparison_join.cpp` — DECIMAL on an inequality join side (TPC-H Q17/Q20, runtime error).** `materialize_expression_join_keys` skipped inequality conditions entirely, leaving both sides to the mixed join's inline cuDF AST predicate. cuDF AST can neither cast to nor compute with DECIMAL, so `cast(l_quantity as decimal(38,5)) < 0.2 * avg(l_quantity)` failed with *"failed to translate mixed join inequality conditions to cuDF AST predicate"*. Inequality sides now go through the same materialization as routed null-safe keys: anything beyond a column reference or an AST-supported cast is projected into a column below the join, leaving a plain column comparison for the predicate. Equality keys are unchanged. Net: +23/−7, no header or public-API change, no behaviour change for plans that already worked (the MARK early-return only fires in the previously-throwing window; materialization only kicks in for expressions the AST could not evaluate). ## Verification - TPC-H SF1, all 22 queries, through the FFI path **and** through the transparent path: results match the DuckDB baseline before/after (the transparent path exercises the unchanged code paths for regression). - Q16 no longer aborts; Q17/Q20 no longer fail at runtime on the FFI path. - No new unit test in this PR: the failing shapes only come out of the Substrait consumer. Happy to add a `test/cpp/exec/test_sirius_ffi_fragment.cpp`-style case if a reviewer wants one — say so and I'll push it. ## Relationship to other PRs Independent of the other two PRs from this branch split. sirius-db#1842 (`experimental/doris`) needs this one at runtime for Q16/Q17/Q20 but does not build against it. --------- Co-authored-by: morningman <moringman@apache.org> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…irius-db#1818) (sirius-db#1843) ## Description Refactors the existing dynamic-filter implementation without adding multi-batch build support. - Moves filter publication out of the hash join into a session that handles completion, cancellation, and cleanup. - Tracks which producers can still publish filters, so consumers can tell when no more filters will arrive. - Gives each filtering pass a consistent snapshot for choosing filters, applying them, and measuring their usefulness. - Keeps filter storage alive while GPU work uses it and adds tests for these lifetimes. Build data is still delivered before filters are constructed and copied to other GPUs. There are no cuCascade changes. No performance comparison is included. ## Checklist - [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist - [x] Cover changes with new or existing tests - [x] Document configuration changes in code and summarize in the description above - [x] Update human and agent documentation (README.md, docs/, skills, CLAUDE.md) ## References Addresses sirius-db#1820 of issue tracker sirius-db#1818
Next towards sirius-db#1734. This one introduces a shared object target. Closes sirius-db#1733.
Fixing one more nightly failure (e.g. https://github.com/sirius-db/sirius/actions/runs/36051990310/job/107809555267).
Our v1.5 [patches](duckdb/duckdb@v1.5-variegata...sirius-db:duckdb:v1.5.5-patches) were [backported](duckdb/duckdb#26102), so we can revert our submodule to upstream DuckDB.
We don't really need to keep and track `duckdb-python` in our repo. The extension we build should just load in a compatible DuckDB's Python package.
Next step in sirius-db#1734. This splits out the DuckDB dep into it's own target. This later allows us to fully own the build i.e. no DuckDB extension build infra in the main project.
This fixes a few more deprecation warnings (that show up in the nightly build).
Follow-up of sirius-db#1870. Replacing `cmake/sirius-sources.cmake` with `CMakeLists.txt` files per subtree, included via `add_subdirectory`.
## Description Large and remote parquet scans need to avoid eagerly reading and decoding data that a query may never consume. This adds frugal caching and dynamic I/O so physical requests can be split, aligned, coalesced, staged, and served incrementally according to each backend's capabilities. The layer includes: - partial-read cache fills with reader/evictor arbitration and explicit lifecycle handling; - reactor-owned request splitting and bulk prepared-slice APIs for REST, io_uring, and kvikIO; - query-event-driven readahead, scheduling, and scan-manager integration; - parquet materialization and device-copy support; - `reset_sirius_cache()` and updated cache/backend configuration; - `pin_table(..., tier='parquet')` for retaining undecoded parquet ranges in the I/O cache; - updated documentation, performance scripts, and I/O benchmarks. Tests cover cache state transitions, request planning, backend behavior, readahead lifecycle, configuration, reset behavior, and parquet scan sizing. ## Validation - `pixi run pre-commit run -a` - `pixi run cmake --build build/release -j 4` (including all standalone I/O benchmark targets) - focused lifecycle/configuration/request tests: 104 test cases, 805 assertions - focused cache tests: 48 test cases, 2,099,946 assertions ## Checklist - [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements. - [x] Cover cache, reactor, readahead, scan-manager, and parquet behavior with tests. - [x] Document backend, cache, and pinning configuration changes in code and user documentation. - [x] Update performance scripts, benchmarks, and agent documentation. ## Stack This was layer 6 of the frugal native-I/O stack; the lower layers (sirius-db#1773, sirius-db#1774, sirius-db#1776, and sirius-db#1772) are already merged. Supersedes sirius-db#1771: `dev` was folded into `main` upstream, and this branch is now rebased directly onto `main`'s tip as a normal, non-stacked PR. ---- GB300 benchmarks: profile lukewarm. hot cold --------------------------- q1 4.0062 2.7908 q2 1.0207 0.3670 q3 1.7495 1.1402 q4 1.9325 0.8274 q5 2.1787 1.1928 q6 0.9368 0.9093 q7 1.3715 1.4290 q8 2.4953 1.1644 q9 2.8493 3.0254 q10 2.5605 1.8375 q11 0.4259 0.3570 q12 1.2622 1.3134 q13 2.1456 0.9090 q14 1.1204 1.1168 q15 1.0552 0.9031 q16 0.4655 0.3556 q17 1.3752 1.2920 q18 2.6438 2.2206 q19 1.4832 1.4426 q20 1.2476 1.2262 q21 2.7803 2.7661 q22 0.3547 0.3392 ------------------------------- Total 37.4606 28.9252 cold Query iter0 ------------------- q1 3.7907 q2 0.9496 q3 2.5600 q4 1.6931 q5 3.0860 q6 2.0780 q7 3.1990 q8 4.3953 q9 5.0900 q10 3.2508 q11 0.6380 q12 2.3308 q13 2.6227 q14 2.3591 q15 2.7282 q16 0.5409 q17 2.4587 q18 3.0535 q19 2.8955 q20 3.5231 q21 4.1316 q22 0.7383 ------------------- Total 58.2130 --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Matthijs Brobbel <m1brobbel@gmail.com>
Bumps CUDA 13 env to latest 13.4 release.
…pare walk (sirius-db#1578) # perf(scan): cache + fuse + parallelize the duckdb-native metadata prepare walk **Branch:** `pr/unpinned-metadata-walk-cache` → base: **`main`**. Rebased onto current `main` (2026-09-29). The two new headers follow sirius-db#1806's move to `src/op/scan/`, and the new sources are registered in `src/op/CMakeLists.txt` and `cmake/sirius-test-sources.cmake` per sirius-db#1905. ## Background: what the "metadata prepare walk" is When Sirius scans a table stored in DuckDB's native format, it first has to plan the scan before touching any data. That planning step walks over every **row group** (DuckDB stores a table as a sequence of row groups, ~122k rows each; a SF1000 `lineitem` has ~49k of them) and, for each one, asks DuckDB for: 1. its geometry (where it starts, how many rows it holds), 2. the min/max statistics of every column that has a pushed-down filter, so whole row groups that can't match can be skipped ("pruning"), 3. the max string length of every projected `VARCHAR` column, to make sure no row group contains an oversized string that the GPU string decoder can't handle (if one does, Sirius refuses the GPU path for that scan). We call this the **metadata prepare walk**. Today it is done from scratch on every query, serially, on the query thread, while the query-lifecycle slot is held. On a `lineitem`-sized table at SF1000 it costs **57–59 ms warm and 1.33 s cold** per query. Since sirius-db#1548 merged, tables whose data is already **pinned** in GPU memory skip this walk entirely. This PR is for everything else: **unpinned** tables, including the first query against a table before its data is pinned. Those still have to do the walk, so this PR makes it cheap. ## What changes Three things, in order of how much of the work they remove: ### 1. Cache the per-table snapshot Steps 1–3 above read data that only changes when the table's physical layout changes (insert, delete, checkpoint). So we now capture it once per table into an immutable **snapshot**: the row-group geometry plus a copy of every row group's column statistics. Later queries against the same table reuse the snapshot instead of asking DuckDB again, which means no storage reads. Getting this right under concurrent writers is where most of the care went: - The capture is taken in a single pass while holding DuckDB's segment-tree lock, so it can't observe a half-committed table. If a commit lands mid-capture anyway, the capture is detected as **torn** (inconsistent) and retried rather than cached. - On every reuse the snapshot is validated against the live table by comparing row-group identities via `weak_ptr`. A `weak_ptr` can only lock back to the exact object it was created from, so a row group that was freed and a new one allocated at the same address is correctly seen as different. (This is the "ABA-safe" property: it protects against the "A was replaced by B, then something else that looks like A" bug.) - If the current transaction has uncommitted appends to the table (DuckDB's `LocalStorage`), the cache is bypassed for that query, because the snapshot can't see those rows. ### 2. Cache the per-query-shape result on top of that Steps 2 and 3 depend not just on the table but on **which columns are projected and which filters are pushed down**. Two queries with the same projected columns and the same filter predicates produce the same pruning decisions and the same oversized-string verdict. So we additionally cache the *outcome* of the walk, keyed by (projected column set, pushed-down filter set). - Filters are compared with `TableFilter::Equals`, restricted to filter types whose comparison actually covers all their parameters, so we never treat two different predicates as equal. - Each cached result is tagged with the snapshot generation it was computed from and is discarded the moment the snapshot rebuilds, so a result can never be served against a table layout it wasn't computed for. A hit here reduces the whole prepare step to a single validity probe, with no statistics checked at all. ### 3. Make the remaining walk faster when it does have to run When neither cache hits (first query against a table, or a new predicate), the walk still has to look at statistics. Previously that was one full pass over all row groups **per filter column** plus one **per varchar column**. Now it is a single fused pass over the row groups that checks all columns at once, and that pass is parallelized across row groups. Parallelizing this is safe because `RowGroup::GetStatistics` takes its own lock per row group and column and returns a self-contained copy, and every `TableFilter::CheckStatistics` implementation is `const` and read-only. Only `GetPartitionStats` has to stay serial (it touches `LocalStorage` and the `ClientContext`). Workers write to disjoint pre-indexed slots, and when a refusal is needed the result is reduced to exactly the same (column, row group) the old serial passes would have reported, so the observable behaviour is unchanged. **Env knobs:** `SIRIUS_METADATA_WALK_THREADS` caps the worker count (`1` = serial); `SIRIUS_DISABLE_NATIVE_METADATA_CACHE` restores the old uncached walk. ## Key review point: why cache invalidation is probe-based rather than counter-based The tempting cheap approach is to key the cache on the table's last commit id (`GetLastCommit`). That is **wrong**: `CHECKPOINT` restructures row groups (compaction) without bumping the commit counter, so a counter-keyed cache would serve stale row-group geometry after a checkpoint. The identity probe described in section 1 is what actually makes this safe. There is an explicit "rebuilds after checkpoint compaction" test so that a future "optimization" back to counter keying fails CI. ## Measured (SF1000 `lineitem`, 48,849 row groups, unpinned, warm; `[walk_bench]` harness) | query shape | before | cache disabled (fused + parallel only) | first query (builds snapshot) | new predicate (snapshot hit) | repeat (result hit) | |---|---|---|---|---|---| | q1-heavy (6 varchars) | 56.6 ms | 19.1 ms | 12.4 ms | 7.6 ms | **6.1 ms (−89%)** | | q1 (2 varchars) | 30.0 ms | 14.8 ms | 8.5 ms | 6.9 ms | 5.5 ms | Column meanings: "cache disabled" isolates section 3 (`SIRIUS_DISABLE_NATIVE_METADATA_CACHE=1`); "first query" is the first walk after a table changes, which pays for capturing the snapshot; "new predicate" is a query with an unseen filter against an unchanged table, where section 3 runs over cached statistics with no storage reads; "repeat" is an already-seen query shape served by section 2. Cold walk: 1327 ms → 12.9 ms. Pin-served benchmark suite: neutral (+0.6%, within noise). This is a latency win for unpinned tables and first queries, not a suite mover. ## Tests `test/cpp/scan/test_duckdb_native_metadata_cache.cpp` (`[duckdb_native_metadata_cache]`, GPU-free), 16 cases: - cached result equals the uncached walk; repeated walks hit; varied filters are served from one snapshot; - result cache hits on repeated query shapes, misses on new predicates, and drops results when the snapshot rebuilds; - invalidation on committed insert; stays valid across committed deletes; rebuilds after checkpoint compaction; bypasses on transaction-local appends; tables are keyed independently; - varchar overflow refusal is preserved, and sub-limit varchars are accepted; - the fused pass reports the same refusal the old column-outer pass did; the parallel walk is deterministic across worker counts; - a torn snapshot is never served under concurrent commits; prepared walks survive concurrent commits and settle exactly. `test/cpp/scan/bench_metadata_walk.cpp` is the hidden `[walk_bench]` harness used for the numbers above. Run on the rebased branch (current `main` base): `[duckdb_native_metadata_cache]` — 16/16 green (1392 assertions); neighbor sweep `[scan],[filter]` — 395/395 green (462,504 assertions). 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…nts (sirius-db#1847) ## Description Compressed-materialization tests currently add six counters, six recording methods, and a snapshot API to `SiriusContext`. This moves those observations onto the query event publisher introduced in sirius-db#1776. Scan, pinning, partition, and planner producers publish an activity and count; only test subscribers aggregate counters. This follows up on the [review discussion in sirius-db#1765](sirius-db#1765 (comment)) and targets `dev` independently of that PR. A subscriber `flush(timeout)` waits for callbacks for all deliveries preceding its captured boundary. Pending event IDs are tracked individually because producers can enqueue out of order. The wait releases the publisher's routing lock and reports timeout, shutdown, or delivery failure instead of allowing tests to assert against partial observations. Existing before/after test helpers subscribe on their first snapshot and flush before reading results, including assertions that no activity occurred. The change preserves each observation's existing units and producer locations. It adds no configuration setting and changes no query execution or partition-sizing decisions. Event tracking and counter aggregation tests belong together here because asynchronous delivery must be synchronized before replacing the existing synchronous counters. ## Checklist - [x] Read CONTRIBUTING.md and keep this independent change scoped for review. - [x] Add event and synchronization tests; migrate existing integration-test observation helpers. - [x] Update compressed-materialization documentation. - [x] No user-facing configuration changes. - [x] Run the full Linux/CUDA build and affected integration tests.
…ltered decode, and key-type generalization (INT8..UINT64, DATE/TIMESTAMP, DECIMAL, STRING, nullable keys) (sirius-db#1557) > **2026-09-29: rebased onto `main` at `b729ac0c` (the repo's default branch was renamed from > `dev`) and review round 2 addressed.** Head `8c131452`; previous tip `b04aba21`. Two commits: > `092db3d9` is the PR itself (rebased; rebase-only conflicts in `CMakeLists.txt` -> main's > per-directory source lists, the publisher's new publication-session `try` block from sirius-db#1843, and > the scan operator's `snapshot.entries()` loop), and `8c131452` holds only the changes made in > response to @joosthooz's seven threads, so the review delta is isolated: > > - **Narrowed build keys are published uniformly** (thread 1, the "silently dropping probes" > one): the DATE-only cast and the `==` zone-map gate are replaced by one domain rule > (`membership_same_family`); zone-map bounds are always restored to the recorded type, and > membership filters are pushed only to bindings whose probe type the domain can read > (`membership_probe_compatible`, counted in `bindings_skipped_incompatible_probe`). New tests: > INT64 key from an INT16 build, DECIMAL64 key from a DECIMAL32 build, compat gate. > - **One source of truth for the carrier rule**: `sirius::value_fits<To>(v)` > (`helper/numeric_carrier_rule.hpp`, `__host__ __device__ constexpr`) is used by the device > `probe_key_convert` and the host `numeric_range_fits`; `integer_storage_type` is the single > timestamp->integer mapping (all units map; the narrowing helper deliberately applies it to DATE > only, documented); `column_values_fit` in `numeric_narrowing` replaces the decimal-specific fit > check; `narrow_domain_of(cudf::data_type)` is the one carrier-family list. > - **Device == host test**: every (rep, carrier) pair at every boundary value vs `value_fits`. > - **Dedup**: `compute_mask` prior-free forwarder defined once in `sirius_mask_applicable`; > constructor prologue shared as `prepare_membership_build`; probe kernel prologue/launch shared > as `membership_probe_functor` / `run_membership_probe`, each filter contributes only its lookup. > > Kernel count unchanged (19 per filter kind). Verified at `8c131452`: release + clang-debug > builds clean; `[dynamic_filter]` 306/306, `[fused_scan_filter]` 27/27, `[numeric_narrowing]` > 14/14; full `make test` 3677/3680 with the 3 failures being a since-fixed assertion in the new > narrowing test (tag re-run green) and two environment/flaky cases unrelated to this PR > (`sirius_config forces the sirius backend for multi-GPU` on a 1-GPU host; `test_explicit_eviction` > from sirius-db#1908 passes in isolation). `pre-commit run -a` clean. > > **2026-09-21: rebased onto `dev` at `b1909fcc` and squashed to one commit** (new head > `b04aba21`; previous tip `743eed39`, incl. the review-cleanup commit, is preserved on the > fork as `pr/carrier-native-df-probes-rebased-dev`'s ancestor and in the reflog). Squashed > because sirius-db#1817 (`rmm::cuda_stream_view` -> `cuda::stream_ref`) rewrote the signatures every one > of the 8 commits touched, so a per-commit rebase conflicted 8 times over. Adaptations to dev: > > - **sirius-db#1806 header split**: `src/include/` is gone; the new files now live beside their sources > (`src/cuda/dynamic_filter_probe.cuh`, `src/op/dynamic_filter/dynamic_filter_key_domain.hpp`) > and the `selection.hpp` edit followed that header to > `src/compression/simpatico_codegen/src/codegen/selection/`. Include paths are unchanged. > - **sirius-db#1817 stream migration**: all PR code and tests use `::cuda::stream_ref` (`.get()`, > `.sync()`); no `cuda_stream_view` remains in PR-touched files. > - `restore_probe_to` call sites that dev still had are gone with it; no references remain. > > Verified at `b04aba21`: `pixi run make` and `pixi run make clang-debug` both clean, > `[dynamic_filter]` 277/277, `[fused_scan_filter]` 27/27, `make test` 3367/3367 on an idle GPU, > `pre-commit run -a` clean. > > **2026-09-17: scope extended.** Six commits (`d219a326`..`b352168c`) generalize the > membership probes to every fixed-width integer, DATE/TIMESTAMP, DECIMAL, STRING and nullable > keys via a key-domain dispatch layer — see **section 3** below. Head moved `80801eb1` -> > `b352168c`; full `make test` 3293/3293. > > **2026-09-16: rebased onto `dev` at `f54e6a10`** (new head `80801eb1`; the previous tip > `30fe3dd9` is preserved on the fork's reflog and referenced here). Still a single commit. > Three files conflicted: > > - `CMakeLists.txt` — test-source list; kept both dev's new iceberg/puffin tests and > `test_fused_membership_mask.cpp`. > - `src/cuda/sirius_dynamic_bloom_filter.cu` — include block only. Dev (sirius-db#1809) dropped the > custom `sirius_bloom_policy` in favour of `cuco::default_filter_policy`, so the > `<cuco/hash_functions.cuh>` include this PR carried is dead and was dropped too; the > `bloom_contains` Bulk kernel is unchanged. > - `src/compression/simpatico_codegen/src/simpatico_codegen.cpp` — dev landed the T9 > positional keep mask (sirius-db#1556) as an extra AND source appended *after* wave 1, while this > PR moves the combine *before* the sequential membership cascade. Resolved by making the > keep mask a cascade prior: it is uploaded on `s0` before the combine and counts as > "another source", so a membership-only chunk *with* a keep mask now takes the sequential > pruned path. With no static range/bool8 strip nothing else writes `combined`, so the keep > mask lands there directly instead of a separate strip. The "every probe declined" fallback > is unchanged and still correct on that path (`keep_mask_applied` stays false, the caller > applies the mask itself). A new case in `test_fused_membership_mask.cpp` covers it. > > The "Deliberately not carried over" paragraph below is therefore partly stale: the T9 > keep-mask **is** now a prior source; pair (col-vs-col) filters remain out. > > Verified on this rebase: `[dynamic_filter]` 241/241, `[fused_scan_filter]` 22/22 (5 fused > membership cases), `make test` 3257/3257 — 36 `gpu_execution iceberg` cases failed on the > first pass only because the `iceberg` DuckDB extension was not provisioned on the box; > all 36 pass after `INSTALL avro; INSTALL iceberg`. `pre-commit run -a` clean. > **Rebased onto `dev`.** This PR previously sat on the closed fused-scan-filter > (sirius-db#1391) + late-mat (sirius-db#1409) stack — the "wave-2 lineage" the old description > referred to. Both were reimplemented and merged as sirius-db#1474 and sirius-db#1524, so the > branch is now a **single commit off `dev`** with no base-branch diff > pollution. The old wave-2 tip is preserved at `a9ac3d88`. ## 1. Carrier-native probes `compute_mask` on all three membership filters required the probe column's type to equal the filter set's type. Compressed materialization stores a bounded `BIGINT`/`INTEGER` key in the narrowest fitting signed carrier, so a decode-time probe sees `INT32` against an `INT64`-built set. The probes now accept any signed-integer carrier (`INT8`..`INT64`) and convert per element in-kernel. A value the key domain cannot represent is a definite non-member, so the no-false-negative contract holds. ### This replaces `restore_probe_to`, and the two were measured against each other sirius-db#1555 fixed the same defect by materializing a widened copy of the probe column before the lookup. The two are **independent fixes for the same problem, derived two weeks apart** — neither the wave-2 base nor T2's own branch had `restore_probe_to`; it appeared during sirius-db#1555's own rebase. Merge order rather than merit picked the winner, so I A/B'd them. Pooled allocator (production hands these calls a pooled `mr`; an unpooled one charges the restore a driver allocation the real path never pays), both arms interleaved in one process, median of 15 warmed reps: | filter | rows | filters/col | A: restore | B: in-kernel | B/A | |---|---|---|---|---|---| | hash IN-list | 1M | 1 | 23.6 µs | 19.0 µs | **0.81** | | hash IN-list | 8M | 1 | 92.1 | 67.7 | **0.74** | | hash IN-list | 8M | 3 | 256.4 | 200.5 | **0.78** | | Bloom | 1M | 1 | 19.6 | 16.2 | **0.82** | | Bloom | 8M | 1 | 67.0 | 46.1 | **0.69** | | Bloom | 8M | 3 | 178.7 | 117.9 | **0.66** | In-kernel is faster at **every point measured**, across two independent runs, and both arms produce identical masks (asserted in the harness). The margin widens with filters per column because `restore_probe_to` sits *inside* `compute_mask` and so re-materializes the same column once per filter — which the post-decode cascade in `dynamic_filter_merge` does routinely, on the default path, not just under the experimental gate. Read the pooled numbers as a **lower bound**: the benchmark pool is fresh, dedicated and uncontended, so production's allocation costs at least this much. Scale: ~3 µs per million rows, so ~18 ms for a full lineitem probe pass at SF1000. Real and consistent, but modest in absolute terms — this is not a headline win, and nothing here recovers the number the old description claimed. `restore_probe_to` had no other callers and is removed with its helpers. ## 2. Mask-aware probing `sirius_mask_applicable` gains a prior-keep-mask `compute_mask` overload (packed 1 bit/row); dead rows skip the set/Bloom lookup. The wave-1 orchestrator runs membership probes sequentially on `s0` after combining the other mask sources, handing the combined words to each probe as its prior and folding each result back in, so later probes see earlier survivors. Membership-only chunks keep the concurrent, prior-free path. The prior is a hint only: ignoring it is sound because the caller ANDs with that same mask. **This half is unmeasured and is the open question on the PR.** It trades concurrency for pruning: a probe that previously ran round-robin on its own pool stream now serializes on `s0` behind the combine, and only the lookup is pruned — the full-width key decode is still paid. At high static keep-rate that may not pay for the lost overlap. **Reasonable to split out** if you'd rather land item 1 on its own measurement and take item 2 separately. ## 3. Key-type generalization: membership probes beyond `INT32`/`INT64` (added 2026-09-17, commits `d219a326`..`b352168c`) Item 1 fixed the *carrier* mismatch for integer keys. This section removes the *type* gate entirely. All three membership filters (hash IN-list, small IN-list, Bloom) previously accepted only `INT32`/`INT64` keys with `null_count()==0`; any other equality key reached the publisher and was silently assigned no membership filter (`choose_membership_filter` -> `none`). Type gating lived in three places, none of them the planner: the filters' `supports()`, the publisher's build-column type-equality check, and the join-edge route gate in the publish plan. ### Dispatch design Two orthogonal, explicitly enumerated axes instead of a `cudf::type_dispatcher` over the (key, probe) type pair (which would instantiate hundreds of kernels): - **Key rep** — the device element type the set / needles / Bloom are built over: `int32_t`, `int64_t`, `uint32_t`, `uint64_t`. New host-only header `src/include/op/dynamic_filter/dynamic_filter_key_domain.hpp`: `membership_key_family` `{signed_int, unsigned_int, date_days, timestamp, decimal, string_hash}`, `classify_membership_key`, `membership_key_supported`, `membership_probe_compatible`, `membership_build_fits_rep`. These replace every `INT32|INT64` literal in the three `supports()`, the publisher, `dynamic_filter_publish_plan.cpp`, and `direct_route_admissible`. - **Probe adapter** — a device functor chosen on the host from `(domain, probe.type())`: `integral_probe_adapter<ProbeRep, KeyRep>` (generalizes item 1's `probe_key_convert`; signedness-aware, range-checked, sentinel-aware), and `string_hash_adapter`. Kernels (`set_contains`, `small_in_list_scan`, `bloom_contains`) take the adapter, a `probe_validity` bitmask reader and the optional prior keep-mask, and emit a **non-nullable** `BOOL8` mask. Instantiation budget: **19 probe kernels per filter kind** (was 8 on dev, 16 for the integer scaffold, +2 for `__int128` decimal carriers, +1 string). Compile time per `.cu` unchanged within noise (~30 s each, measured cache-busted). ### What each family does | Family | Rep | Build side | Probe side | Notes | |---|---|---|---|---| | `INT8`/`INT16`/`INT32`/`INT64`, `UINT8`..`UINT64` | i32/i64/u32/u64 | any signed/unsigned integer column; the publisher now **accepts a build column arriving at a narrower carrier** when `can_restore_to` holds (previously the key was skipped, `keys_skipped_type_mismatch`) | any carrier of the same signedness family, converted in-kernel; out-of-domain value = definite non-member | sentinel is `min()` for signed, `max()` for unsigned reps | | `TIMESTAMP_DAYS` (DATE), `TIMESTAMP_{S,MS,US,NS}` | i32 / i64 | reads the underlying rep; a DATE build column arriving as `INT8`/`INT16` is **restored to `TIMESTAMP_DAYS`** in the publisher (a carrier-typed set would decline the native probe the post-decode cascade emits) | same unit, or `INT8`/`INT16`/`INT32` carriers for DATE; mixed units decline (unreachable anyway: DuckDB inserts a cast and a cast blocks the scan route) | zone maps already supported these and serve as the test oracle | | `DECIMAL32`/`DECIMAL64`/`DECIMAL128` | i32 / i64 / i64-if-fits | `DECIMAL128` builds are range-probed (`compute_exact_numeric_range`); if the unscaled range fits `int64` the set is `int64`, else **all three membership filters decline** (zone map still publishes) | same scale required; `DECIMAL32`/`64`/`128` carriers all accepted, `__int128` probes range-checked into `int64` | no `__int128` Bloom (deliberately) | | `STRING` | u64 | `cudf::hashing::xxhash_64` fingerprints materialized once per build | `XXHash_64<string_view>` computed **in-kernel** over a `column_device_view`, no materialized copy; `DICTIONARY32` declines (the fused decode reconstructs dictionary/str_split carriers to `STRING` first) | both IN-lists become **no-false-negatives** for strings (fingerprint collisions keep rows); the sentinel `UINT64_MAX` needs no remap — cuco's insert of its empty key is a no-op and the probe keeps rows whose fingerprint equals it | | **nullable keys** (any family) | — | `supports()` no longer requires `null_count()==0`; null build keys are compacted out (`drop_nulls`, as Bloom already did); the small IN-list tier gates on *valid* rows; the publisher sizes/tiers on `valid_rows` | null probe rows emit `false` via the `probe_validity` prologue; `copy_bitmask` tails removed | exact, not approximate: null-safe keys are never admitted and the join runs `null_equality::UNEQUAL`, so a NULL never matches on either side | Not done, on purpose: `FLOAT32/64` (no benchmark uses float keys; would need -0.0/NaN canonicalization), `BOOL8`, `LIST`/`STRUCT`, `HUGEINT` (its lossy `INT64` mapping is a separate bug). ### Where the expected TPC-H wins turned out not to exist TPC-H has **only `INTEGER`/`BIGINT` join keys**, so this section is inert on TPC-H beyond the narrow-build-carrier acceptance and nullable-key handling. In particular: - q2's `ps_supplycost = min(...)` is planned as a `FILTER` above a `DELIM_JOIN`, not as a join condition — there is no DECIMAL join key to admit. - q15's `total_revenue = (subquery)` *is* a DECIMAL128 join condition, but its build side is `AGGREGATE max(...)` over a `CTE_SCAN`, which `build_relation_is_opaque` / `build_subtree_is_filtering` do not recognize as evidence. That is the evidence policy, upstream of admission, and is untouched here. The beneficiaries are TPC-DS. A survey at SF10 on the integer-only scaffold found that of the 21 queries with STRING/DATE/DECIMAL/nullable join keys, only ~7 currently run on the GPU at all (the rest fall back on UNION / WINDOW / DISTINCT / mixed-inequality joins, or crash: q10 aborts on a MARK-join mode assertion). Those seven *do* show the key being declined today: q83 (`d_date` -> `none`), q59 (`s_store_id` -> `none`), q30/q81/q64/q69 (nullable INT keys downgraded IN-list -> Bloom). No TPC-DS timing is claimed in this PR. **SF1000 TPC-H QphH, same recipe as August, no patched libcudf:** dev `f54e6a10` 7.79M (Power 10.15M / Tput 5.98M); this branch at `80801eb1` (items 1+2 only) median of 3 clean runs 7.95M (Power 10.45M / Tput 6.04M) — Power +3.3%, inside the ±4.5% run-to-run band. Firmer than the score: dev hit GPU-OOM retry-cap CPU fallbacks in 3 of 4 throughput phases, this branch in 0 of 3. Per-query the membership-filter consumers moved most (q17 −11%, q8 −9%, q7/q20 −7%, q19 −6%). The type-generalization commits were not scored separately: they cannot change any TPC-H plan. **Re-measured 2026-09-21 at the rebased head `b04aba21` vs dev `b1909fcc`** (4 clean interleaved runs each, medians): Power 10.49M vs 10.19M (+2.9%), Throughput 6.02M vs 5.99M (+0.5%), QphH **7.94M vs 7.83M (+1.5%)**, stream-0 7.85 s vs 8.07 s. All four PR Power values sit above all four dev values; q19 −17%, q17 −11%, q8 −7%, q20 −6%, q7 −5%; nothing slower than dev by >0.3%. Side finding, not caused by this PR: on this dev base the 7-stream throughput phase hits GPU-OOM retry-cap CPU fallbacks in 6 of 10 dev attempts (2 of 6 with the PR; 0 of 1 on the f54e6a1-based control), which then either kills the streams on DuckDB's 32 GB limit or spills hundreds of GB to disk. Power phases are unaffected. ### Tests for this section `[dynamic_filter]` 277 cases / `[fused_scan_filter]` 27 cases; full `make test` 3293/3293 on an idle GPU at `b352168c`. New coverage, per family and for all three filters, with and without a prior mask, against a host oracle and (for DATE/DECIMAL) the lowered zone-map AST as an independent oracle: cross-carrier identity `INT8..INT64` x `{INT32,INT64}` keys vs. the old path; `INT8`/`INT16`/`UINT*` keys and unsigned sentinel semantics; `TIMESTAMP_DAYS` x `{DAYS, INT8, INT16, INT32}` probes and unit-mismatch declines; `DECIMAL{32,64,128}` x `DECIMAL{32,64,128}` probes, scale mismatch declines, DECIMAL128 fit -> int64 set vs. overflow -> decline; STRING host XXH64 reference pinned against `cudf::hashing::xxhash_64` over a corpus covering every XXH64 tail path (empty, 1 KiB, UTF-8, embedded NUL); nullable builds x nullable probes x 4 carriers, sliced-probe offsets, all-null build/probe; publisher accepting `INT8`/`INT16` build carriers, `DECIMAL32`-carrier for a `DECIMAL64` key, `INT16`-carrier DATE; fused decode of `INT8`-carrier, `INT16`-carrier-DATE, `DECIMAL32`-carrier and dictionary-`STRING` chunks against the widened sets. `docs/super-sirius/dynamic-filters.md` gains a "Key types" section with this table. ## The original's numbers do not transfer The old description claimed SF1000 Power **+7.2%**, q19 **−52%**. Those were measured against a stack where a declining probe *threw* and failed the entire filtered decode. sirius-db#1524 changed that to a stand-down to the AND identity, and sirius-db#1555 removed the decline for narrowed carriers altogether. **Treat the old numbers as void.** ## Notes - **The multi-probe cascade is unreachable in a default build.** `SIRIUS_EXP_FUSED_SCAN_MAX_MEMBER` defaults to `1`, so the fold-back never executes; the new test arms it to 2. Worth deciding whether that default moves. - **Small-in-list was not benchmarked** — the A/B covered hash IN-list and Bloom. - Item 2's decode-time path sits behind `SIRIUS_EXP_FUSED_SCAN_FILTER` (off unless set). Item 1 is on the default path via the post-decode cascade. ## Deliberately not carried over `dev`'s reimplementation dropped pair (col-vs-col) filters and the T9 visibility keep-mask (still open as sirius-db#1556), both of which the original also primed priors from. Priors here are ranges and dictionary-answered equalities only. A declining probe in the sequential cascade leaves the mask untouched, matching `dev`'s all-ones stand-down rather than the original's throw. ## Related, not fixed here The build side has the mirror defect with a worse failure mode: `dynamic_filter_publisher.cpp` compares the runtime build column against the plan's recorded logical type and **skips the key entirely** on a mismatch (`keys_skipped_type_mismatch`), so a narrowed build key publishes no filter at all — probe-side tolerance cannot help there. Already instrumented via `dynamic_filter_stats`; worth checking whether that counter fires. ## Tests - `test/cpp/operator/test_dynamic_filter_probe.cpp` (`[dynamic_filter][probe]`, 8 cases): heterogeneous carriers, decimal/date/float refusal, sentinel conservation, prior-mask all-dead/all-live/patterned, null propagation. - `test/cpp/scan/test_fused_membership_mask.cpp` (`[fused_scan_filter]`, 4 cases): prior handed alongside a static range, prior-free membership-only, a two-probe cascade proving the fold-back, all-dead prior, INT32 carrier vs INT64 set. Against current `dev`: `[probe]` 8/8, `[fused_scan_filter]` 4/4, `[dynamic_filter]` 240/240, **full suite 3035/3035** (32,788,590 assertions). `pre-commit` clean. The A/B harness was scaffolding and is not in the diff. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…arrocks (sirius-db#1911) Bumps [thiserror](https://github.com/dtolnay/thiserror) from 2.0.20 to 2.0.21. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/dtolnay/thiserror/releases">thiserror's releases</a>.</em></p> <blockquote> <h2>2.0.21</h2> <ul> <li>Fix parsing of generic unit variants in display expressions (<a href="https://redirect.github.com/dtolnay/thiserror/issues/459">#459</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/dtolnay/thiserror/commit/b1827ee06f81a7d7f954e676771e19a97f9f8e3b"><code>b1827ee</code></a> Release 2.0.21</li> <li><a href="https://github.com/dtolnay/thiserror/commit/58037b575a2a5543d0d54a49360554325b49f55a"><code>58037b5</code></a> Merge pull request <a href="https://redirect.github.com/dtolnay/thiserror/issues/459">#459</a> from dtolnay/turbofish</li> <li><a href="https://github.com/dtolnay/thiserror/commit/f82a0cf8f06c71263c2caa26e759f337d2778153"><code>f82a0cf</code></a> Keep track of nested turbofish depth</li> <li><a href="https://github.com/dtolnay/thiserror/commit/72ea49262d6412ccbc08f1696dda8626409e657d"><code>72ea492</code></a> Raise required compiler to Rust 1.77</li> <li><a href="https://github.com/dtolnay/thiserror/commit/72eea0d4ddb17ff9fa873aa74469de850021b22e"><code>72eea0d</code></a> Resolve io_other_error clippy lint in tests</li> <li><a href="https://github.com/dtolnay/thiserror/commit/07f09a2ec934df58508ce4fa5e7fbf27cca552c2"><code>07f09a2</code></a> Raise required compiler to Rust 1.74</li> <li><a href="https://github.com/dtolnay/thiserror/commit/2715388e8cc5787b3643251f7a7a637e89376688"><code>2715388</code></a> Update ui test suite to nightly-2026-09-22</li> <li><a href="https://github.com/dtolnay/thiserror/commit/5a306c7d0a8588caaaaa6a7567aeb25c1c10719b"><code>5a306c7</code></a> Update ui test suite to nightly-2026-09-05</li> <li><a href="https://github.com/dtolnay/thiserror/commit/ef9383b37d7c96b01e10df35d3e9e313c7b98b5a"><code>ef9383b</code></a> Update ui test suite to nightly-2026-08-22</li> <li><a href="https://github.com/dtolnay/thiserror/commit/8336b8407fbfe69177c504cdbc379db20cf6f131"><code>8336b84</code></a> Update ui tests for version 2.0.20</li> <li>See full diff in <a href="https://github.com/dtolnay/thiserror/compare/2.0.20...2.0.21">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…an_manager_state) (sirius-db#1366) This is PR 3 from a stack of PRs (it goes on top of sirius-db#1364 and sirius-db#1365) which together will make Sirius internals be able to handle Concurrent query execution. There is more to come. ## Summary This PR addresses the scan_manager. It primarily takes all its internal structures that are query specific and puts them into a query_scan_manager_state. Then the scan_manager can hold a map of these states. ## Background The scan manager used to hold one query's worth of state in bare members (_providers_by_op, _scan_op_order, _metadata_processor, _dispatcher, _pending_mvcc_mask_jobs, _pending_insert_delta_jobs). reset() wiped all of them. This was fine for a single serial query but breaks under concurrent queries ## What changed ### 1. query_scan_manager_state All per-query state is now grouped in a new inner struct, query_scan_manager_state, held as a shared_ptr in a map keyed by query_id: std::map<query_id_t, std::shared_ptr<query_scan_manager_state>> _query_states; mutable std::mutex _query_states_mutex; State is published to the map AFTER prepare_for_query fully builds it, so a failed prepare never leaves a half-built entry visible. The shared_ptr is resolved under the lock and used outside it, so an erase racing a reader cannot pull the state out from under the reader. ### 2. reset(query_id) / reset_all() reset() now takes a query_id and only tears down that query's dispatcher, coalescer, and providers. Other queries' state is not touched. reset_all() is the teardown-only path used by stop() and the failed-query backstop. sirius_context::run_mandatory_cleanup now calls reset(query_id) instead of reset(). sirius_context::drop_query_runtime_state_best_effort (the error backstop) now also calls reset(query_id): prepare_for_query no longer performs a global reset. ### 3. Shared-ownership pin table _pinned_entries changed from: std::unordered_map<std::string, pinned_entry> to: std::unordered_map<std::string, std::shared_ptr<pinned_entry>> A query that matched a pinned entry now co-owns it for its whole duration through the shared_ptr the cached_databatch_provider holds. A concurrent UNPIN on another connection only drops the map slot; the data stays alive until the last serving provider releases it. find_pinned_entry_for_duckdb_table changed from returning a raw pinned_entry const* to a shared_ptr<const pinned_entry> (owning), so plan-time MVCC guards that read the entry hold it alive across the check. make_provider_for_pinned_entry signature changed from pinned_entry const& to std::shared_ptr<pinned_entry const> for the same reason. All pin-table mutations (insert, attach_mvcc_metadata, remove, visit_pinned_entries) are now guarded by _pinned_entries_mutex. try_match_cached_entry snapshots the pin table under _pinned_entries_mutex and matches outside the lock: the match path (can_serve_with_columns, validate, zone-map plan building, MVCC branch) is too slow to hold the lock across. ### 4. Thread pool sizing The pool is now sized num_threads + k_max_concurrent_queries (was num_threads + 1). Each query's coalescer parks exactly one sequencer task in queue.wait_dequeue, only unblocked by that query's own split_provider tasks running on the same pool. With Q concurrent queries, Q threads are parked in sequencers at all times, so the old +1 sizing would deadlock at 2 concurrent queries. k_max_concurrent_queries is left at 1 (matching the old behavior for single-query runs) with a TODO to promote it to a real config option. ### 5. Known gap documented The prefetch cache epoch is a global counter (prefetching_cache::_ticker), so a second query starting bumps the epoch and demotes the first query's prefetched-but-unconsumed chunks to tier-0 in the eviction order. Performance only (not correctness: a live reader holds pin > 0, blocking mark_evicting). The fix belongs in prefetching_cache, not here. A comment now calls this out clearly. ## Changed interfaces sirius_scan_manager::reset() -> reset(query_id) + reset_all() sirius_scan_manager::find_pinned_entry_for_duckdb_table() raw pointer -> shared_ptr (owning) make_provider_for_pinned_entry() pinned_entry const& -> shared_ptr SiriusContext::run_mandatory_cleanup scan_manager_->reset() -> reset(query_id) drop_query_runtime_state_best_effort: scan_manager_->reset(query_id) added ## New test: test_scan_manager_query_state.cpp - two queries register independently and reset drops only one Pins that tearing A's dispatcher down does not kill B's sequencer, and B's split connector still yields splits after A resets. - concurrent queries do not collide on operator id Both queries get operator id 0. Asserts each operator drains splits for its own file only. - reset is a no-op for unknown and already-reset queries The failure backstop calls reset(query_id) unconditionally on every code path; an unknown id must not throw. - reset_all drops every query and stop() still returns A sequencer parked on a dequeue with nothing to feed it would hang stop() indefinitely. - a query with no GPU scan operators registers nothing The matching reset is a harmless no-op rather than touching the map. ## Files changed src/include/scan_manager/sirius_scan_manager.hpp main API changes src/scan_manager/sirius_scan_manager.cpp implementation src/sirius_context.cpp per-query reset calls src/planner/sirius_plan_get.cpp owning shared_ptr test/cpp/scan_manager/test_scan_manager_query_state.cpp (new) test/cpp/scan_manager/test_cached_serving_hardening.cpp shared_ptr entry test/cpp/scan_manager/test_s3_routing_cutover.cpp distinct query ids CMakeLists.txt
# Conflicts: # CMakeLists.txt # pixi.lock # src/creator/task_creator.cpp # src/creator/task_creator.hpp # src/cuda/tae/tae_decode_kernels.hpp # src/cudf/cudf_compat.hpp # src/data/host_tae_representation.hpp # src/data/host_tae_representation_converters.hpp # src/data/sirius_converter_registry.hpp # src/embedding/buffer_budget.hpp # src/embedding/control.hpp # src/embedding/execution_interrupted.hpp # src/embedding/execution_stats.hpp # src/embedding/input.hpp # src/embedding/native_gpu.hpp # src/embedding/plan_bindings.hpp # src/embedding/result.hpp # src/embedding/result_codec.hpp # src/embedding/result_gpu.hpp # src/embedding/source_wakeup.hpp # src/embedding/tae_demand.hpp # src/embedding/tae_read.hpp # src/embedding/terminal_admission.hpp # src/execution/sirius_execution_evidence.hpp # src/include/op/scan/gpu_ingestible_types.hpp # src/op/scan/tae_gpu_ingestible.hpp # src/op/scan/tae_scan_plan.hpp # src/pipeline/gpu_pipeline_executor.cpp # src/pipeline/gpu_pipeline_task.cpp # src/pipeline/gpu_stream_quiescence_error.hpp # src/pipeline/sirius_pipeline.cpp # src/planner/sirius_physical_plan_generator.cpp # src/sirius_c.h # src/sirius_ffi.cpp # src/tae/tae_format.hpp
…-dev-merge # Conflicts: # pixi.lock # src/scan_manager/sirius_scan_manager.cpp # src/scan_manager/sirius_scan_manager.hpp
aunjgr
marked this pull request as ready for review
September 29, 2026 10:53
aunjgr
merged commit Sep 29, 2026
e2e2f08
into
matrixorigin:upstream-dev-merge
14 of 21 checks passed
4 of 7 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Synchronize MatrixOrigin's
upstream-dev-mergeintegration branch withsirius-db/sirius:mainthrough803059a0(DuckDB v1.5.6), retaining the native embedding API needed by MatrixOne. The upstream CMake/source relocation is reconciled with the embedding targets and tests; its per-query scan-manager state is reconciled with bounded per-scan metadata producers. Teardown stops producers, closes full mailboxes/connectors, joins coalescer tasks, then joins producers before destroying callback state.The
moPixi build profile disables C/C++/CUDA compiler caching. A cache hit had supplied dependency files pointing into an old sidecar checkout, allowing mixed-revision objects and invalid test results; a clean MO-owned checkout build now has dependency paths only undermatrixone/third_party/sirius. Upstream main advanced further during this clean validation; those later commits are outside this reviewed snapshot.This does not enable direct-TAE reads from MatrixOne. MO remains reader-only for this milestone, and its draft PR will pin the merged Sirius SHA after this PR lands.
Validation in
matrixone/third_party/sirius, usingpixi run --frozen -e moon one RTX 3070: clean no-cache SDK/C-smoke build and execution; all 3,627 main unit cases (39,994,537 assertions); 76 native control cases; 5 native binding cases; 11 native GPU cases; two-stream native TAE/result integration; eight SDK-exporter tests; complete pre-commit suite. DuckDB 1.5.6's Iceberg extension was installed for the full unit run. No container or sidecar service was used. Combined MO binary and public SQL/TPCH validation remain separate gates of MatrixOne's draft bridge and later MO-reader PRs.Checklist
References
Refs matrixorigin/matrixone#28966.