Skip to content

build: sync embedding branch with upstream Sirius main - #23

Merged
aunjgr merged 75 commits into
matrixorigin:upstream-dev-mergefrom
aunjgr:feature/28966-sync-upstream-main
Sep 29, 2026
Merged

aunjgr merged 75 commits into
matrixorigin:upstream-dev-mergefrom
aunjgr:feature/28966-sync-upstream-main

Conversation

@aunjgr

@aunjgr aunjgr commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Synchronize MatrixOrigin's upstream-dev-merge integration branch with sirius-db/sirius:main through 803059a0 (DuckDB v1.5.6), retaining the native embedding API needed by MatrixOne. The upstream CMake/source relocation is reconciled with the embedding targets and tests; its per-query scan-manager state is reconciled with bounded per-scan metadata producers. Teardown stops producers, closes full mailboxes/connectors, joins coalescer tasks, then joins producers before destroying callback state.

The mo Pixi build profile disables C/C++/CUDA compiler caching. A cache hit had supplied dependency files pointing into an old sidecar checkout, allowing mixed-revision objects and invalid test results; a clean MO-owned checkout build now has dependency paths only under matrixone/third_party/sirius. Upstream main advanced further during this clean validation; those later commits are outside this reviewed snapshot.

This does not enable direct-TAE reads from MatrixOne. MO remains reader-only for this milestone, and its draft PR will pin the merged Sirius SHA after this PR lands.

Validation in matrixone/third_party/sirius, using pixi run --frozen -e mo on one RTX 3070: clean no-cache SDK/C-smoke build and execution; all 3,627 main unit cases (39,994,537 assertions); 76 native control cases; 5 native binding cases; 11 native GPU cases; two-stream native TAE/result integration; eight SDK-exporter tests; complete pre-commit suite. DuckDB 1.5.6's Iceberg extension was installed for the full unit run. No container or sidecar service was used. Combined MO binary and public SQL/TPCH validation remain separate gates of MatrixOne's draft bridge and later MO-reader PRs.

Checklist

  • Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
  • Cover changes with new or existing tests
  • Document configuration changes in code and summarize in the description above
  • Update human and agent documentation (README.md, docs/, skills, CLAUDE.md)

References

Refs matrixorigin/matrixone#28966.

dependabot Bot and others added 30 commits September 15, 2026 19:41
Bumps [cxx](https://github.com/dtolnay/cxx) from 1.0.200 to 1.0.202.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/dtolnay/cxx/releases">cxx's
releases</a>.</em></p>
<blockquote>
<h2>1.0.202</h2>
<ul>
<li>Support build environments that limit build-script filesystem writes
outside of OUT_DIR (<a
href="https://redirect.github.com/dtolnay/cxx/issues/1760">#1760</a>,
thanks <a
href="https://github.com/gudcks0305"><code>@​gudcks0305</code></a>)</li>
</ul>
<h2>1.0.201</h2>
<ul>
<li>Specialize <code>enable_borrowed_range</code> and
<code>enable_view</code> for <code>rust::Slice&lt;T&gt;</code> (<a
href="https://redirect.github.com/dtolnay/cxx/issues/1758">#1758</a>,
thanks <a
href="https://github.com/MixusMinimax"><code>@​MixusMinimax</code></a>)</li>
<li>Register bazel toolchains as dev dependencies (<a
href="https://redirect.github.com/dtolnay/cxx/issues/1759">#1759</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/dtolnay/cxx/commit/983982c5fd5584a5391ae46877e952118e431f73"><code>983982c</code></a>
Release 1.0.202</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/552330e2a705f1b0ed03454eb2fc9b4d702ff556"><code>552330e</code></a>
Touch up PR 1760</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/3a9d010450c5ee145a9394ced0b83a294215a91e"><code>3a9d010</code></a>
Merge pull request 1760 from
gudcks0305/codex/fix-shared-header-fallback</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/97c1cfffcbc51bea167b927dd4acdba147ef3bbf"><code>97c1cff</code></a>
fix: tolerate unavailable shared header directory</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/982aa18796aef5ffea9c26466bd21d8397002cb9"><code>982aa18</code></a>
Release 1.0.201</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/a543f3b5c859b6a11757a4899a0dd367fb2a7a37"><code>a543f3b</code></a>
Lockfile update</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/a237c373198ff5cfe4aa8dc9e2e49ac75643b114"><code>a237c37</code></a>
Reformat with clang-format 22</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/eeabfb536ec143ff52bd1c54a2bc9eecba6881f2"><code>eeabfb5</code></a>
Move static assertions from header to source file</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/47d9bf7a5fae6657712e1e1373f1758b40fca660"><code>47d9bf7</code></a>
Merge pull request <a
href="https://redirect.github.com/dtolnay/cxx/issues/1758">#1758</a>
from MixusMinimax/slice-impl-borrowed-range-and-view</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/5a9c2ba4cebbc3ac62f20175e1bb60c26b82affb"><code>5a9c2ba</code></a>
Merge pull request <a
href="https://redirect.github.com/dtolnay/cxx/issues/1759">#1759</a>
from dtolnay/bazel</li>
<li>Additional commits viewable in <a
href="https://github.com/dtolnay/cxx/compare/1.0.200...1.0.202">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=cxx&package-manager=cargo&previous-version=1.0.200&new-version=1.0.202)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…cks (sirius-db#1800)

Bumps [cxx](https://github.com/dtolnay/cxx) from 1.0.200 to 1.0.202.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/dtolnay/cxx/releases">cxx's
releases</a>.</em></p>
<blockquote>
<h2>1.0.202</h2>
<ul>
<li>Support build environments that limit build-script filesystem writes
outside of OUT_DIR (<a
href="https://redirect.github.com/dtolnay/cxx/issues/1760">#1760</a>,
thanks <a
href="https://github.com/gudcks0305"><code>@​gudcks0305</code></a>)</li>
</ul>
<h2>1.0.201</h2>
<ul>
<li>Specialize <code>enable_borrowed_range</code> and
<code>enable_view</code> for <code>rust::Slice&lt;T&gt;</code> (<a
href="https://redirect.github.com/dtolnay/cxx/issues/1758">#1758</a>,
thanks <a
href="https://github.com/MixusMinimax"><code>@​MixusMinimax</code></a>)</li>
<li>Register bazel toolchains as dev dependencies (<a
href="https://redirect.github.com/dtolnay/cxx/issues/1759">#1759</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/dtolnay/cxx/commit/983982c5fd5584a5391ae46877e952118e431f73"><code>983982c</code></a>
Release 1.0.202</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/552330e2a705f1b0ed03454eb2fc9b4d702ff556"><code>552330e</code></a>
Touch up PR 1760</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/3a9d010450c5ee145a9394ced0b83a294215a91e"><code>3a9d010</code></a>
Merge pull request 1760 from
gudcks0305/codex/fix-shared-header-fallback</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/97c1cfffcbc51bea167b927dd4acdba147ef3bbf"><code>97c1cff</code></a>
fix: tolerate unavailable shared header directory</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/982aa18796aef5ffea9c26466bd21d8397002cb9"><code>982aa18</code></a>
Release 1.0.201</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/a543f3b5c859b6a11757a4899a0dd367fb2a7a37"><code>a543f3b</code></a>
Lockfile update</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/a237c373198ff5cfe4aa8dc9e2e49ac75643b114"><code>a237c37</code></a>
Reformat with clang-format 22</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/eeabfb536ec143ff52bd1c54a2bc9eecba6881f2"><code>eeabfb5</code></a>
Move static assertions from header to source file</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/47d9bf7a5fae6657712e1e1373f1758b40fca660"><code>47d9bf7</code></a>
Merge pull request <a
href="https://redirect.github.com/dtolnay/cxx/issues/1758">#1758</a>
from MixusMinimax/slice-impl-borrowed-range-and-view</li>
<li><a
href="https://github.com/dtolnay/cxx/commit/5a9c2ba4cebbc3ac62f20175e1bb60c26b82affb"><code>5a9c2ba</code></a>
Merge pull request <a
href="https://redirect.github.com/dtolnay/cxx/issues/1759">#1759</a>
from dtolnay/bazel</li>
<li>Additional commits viewable in <a
href="https://github.com/dtolnay/cxx/compare/1.0.200...1.0.202">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=cxx&package-manager=cargo&previous-version=1.0.200&new-version=1.0.202)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…ge extraction (sirius-db#1577)

# perf(scan): see through DATE→TIMESTAMP casts in scan filter range
extraction

**Branch:** `pr/cast-through-range-extraction` → base: **`dev`**. Two
commits, no longer stacked: the change, and the review follow-up for
@mbrobbel's ±infinity / overflow comment. Rebased onto `dev` @
`05148c78` (2026-09-15).

> [!NOTE]
> **Previously stacked on sirius-db#1555**, which has since merged (`3ac080d8`).
This PR was rebased onto `dev`: the ~70 never-merged commits of the
sirius-db#1391 fused scan-filter lineage are gone, and the one remaining change
was **ported** rather than replayed — `dev` had meanwhile refactored the
extraction machinery out of `scan_utils.cpp` into
`op/scan/scan_filter_analysis.{hpp,cpp}`
(`extract_numeric_range_pushdown` → `analyze_scan_filters`,
`decoded_bound{floor,ceil}` → `sirius::numeric_range{minimum,maximum}`).
The diff below is now exactly this change: **5 files, +889 / −24**.

## Problem

DuckDB constant-folds qgen-style date arithmetic (`l_shipdate <= DATE
'1998-12-01' - INTERVAL '72' DAY`) into `CAST(l_shipdate AS TIMESTAMP)
<= TIMESTAMP '1998-09-20 00:00:00'`, which arrives at the scan as an
`EXPRESSION_FILTER`. `analyze_scan_filters` only understood
`CONSTANT_COMPARISON` / `CONJUNCTION_AND`, so `fold_numeric_conjunct`
refused the clause and the whole scan lost its in-decode masking:
full-width decode + per-row cast AST filter + gather-compact. q1 varied
1.288 s vs 0.831 s with the equivalent DATE literal; the same shape
appears in 9 predicates across q1/q4/q5/q6/q10/q12/q14/q15/q20 under
spec-official (qgen) parameters.

## Change

Three additions in `src/op/scan/scan_filter_analysis.cpp`:

- **`to_decoded_bound`**: accepts finite `TIMESTAMP[_S|_MS|_NS]`
constants against DATE columns, lowering ticks / ticks-per-day into the
stored-day domain by floor/ceil (templated on the finer-scale DECIMAL
branch). The bound is EXACT for every comparison op — midnight(d) is
strictly monotonic in d — so midnight and non-midnight constants both
keep full coverage and the residual filter drops; `= non-midnight` folds
to a provably empty range.
- **`fold_comparison_bound`**: the comparison switch lifted out of
`fold_numeric_conjunct` so both filter shapes share one implementation.
- **`fold_expression_conjunct`** + a new `EXPRESSION_FILTER` case,
matching `cmp(CAST(bound_ref AS TIMESTAMP*), ts_const)` in both operand
orders (comparison flipped on swap), recursing through AND conjunctions.

**Edges of the DATE domain** (review follow-up, see
`lower_timestamp_to_days`): `DATE ±infinity` rows compare correctly
because ±INT32_MAX days and ±INT64_MAX ticks are the extremes of their
domains, and **±infinity constants now lower exactly** to the infinite
date's stored days instead of being refused. A finite date whose
midnight overflows the target (~±292,000 years for `TIMESTAMP/_S/_MS`,
~±292 years for `TIMESTAMP_NS`) makes DuckDB raise; a decode range
cannot raise, and neither does the residual `cudf::cast` it replaces, so
the range is deliberately not clipped to the castable window and such
rows are kept or dropped by the instant they denote. The only divergence
from DuckDB, and only on queries DuckDB refuses to answer.

**Refused (coverage cleared, residual kept):** `TRY_CAST`
(NULL-on-overflow diverges from day math on >-side bounds),
`TIMESTAMP_TZ` (midnight is session-timezone dependent), `INT64_MIN`
ticks (below -infinity, never a real instant), and every non-comparison
shape (including the `BETWEEN` DuckDB builds from two cast comparisons
on one column; follow-up).

## Measured (SF1000, GB300, spec-official qgen parameters)

- **Power@Size +9.6% → 10,079,305**
- Power-run suite time **−8.3%**
- q1/q12/q14/q15 each **−36..39%** (the queries whose only selective
predicate is a folded date cast)

> [!WARNING]
> These numbers were measured on the **pre-rebase fused-scan-filter
stack**, not on this branch. They have **not** been re-measured against
current `dev`, whose scan/decode path has since changed (sirius-db#1524 late
materialization, sirius-db#1555, and others). Treat them as the original
motivation, not as a claim about this branch's delta. Correctness on
current `dev` is verified below.

## Amplification history (honest note)

Multiplying the number of fused scans exposed two classes of
pre-existing bugs downstream of scan-produced geometry — neither
introduced by this change, both diagnosed during its qualification:

- two pre-existing Sirius GPU use-after-free races (batch read-lock
lifetime, error-path teardown);
- a pre-existing **CCCL 3.4.0** bug: the warp-specialized (TMA)
`cub::DeviceScan` on Blackwell performs an out-of-bounds `__shared__`
read that silently corrupts scan outputs under concurrent kernel
diversity (memcheck-reproducible with a 22-line bare `ExclusiveSum`;
CTK-bundled CCCL 3.2.0 is clean).

Both hardenings ship separately in **sirius-db#1580** (which supersedes the
now-closed sirius-db#1546).

## Tests

- `test/cpp/scan/test_scan_filter_cast_ranges.cpp` (`[range_pushdown]`):
exhaustive boundary matrix — all 5 ops × midnight / one-µs-either-side /
noon constants, pre-epoch floor, operand flips, AND shapes, S/MS/NS
flavors, extreme finite ticks, ±infinity constants (all ops, both
directions, both operand orders, every flavor), `INT64_MIN` refusal,
bounds not clipped to the castable window, and the refusal set.
- `test/cpp/integration/test_gpu_execution_cast_date_predicates.cpp`
(`[cast_date]`): GPU-vs-CPU equivalence for the folded SQL shapes on
three scan paths. **On current `dev` the fold's ranges only take effect
on a GPU-pinned Parquet entry decoded through the compressed path** (the
DuckDB-native ingestible does not analyze its filter), so besides the
plain and DuckDB-pinned scans the fixture pins bitpacked Parquet twins,
proves via census that the pin compressed, and lifts the fused-decode
selectivity caps so the in-decode mask answers. New tables: `DATE
±infinity` rows (finite constants on every path, infinite constants on
the fold path) and a year-300000 date that raises on CPU while the fold
answers by instant.

Run on this branch, rebased onto `dev` @ `05148c78`:

| Filter | Result |
| --- | --- |
| `[range_pushdown],[cast_date],[fused_scan_filter]` | 17/17 green —
9,788 assertions |
| `[scan],[filter]` neighbor sweep | 374/374 green — 461,053 assertions
|

`pixi run pre-commit run -a` clean on the changed files.

> [!NOTE]
> Found while making the tests bite, pre-existing and not touched here:
the residual `cudf::cast` path multiplies `±INT32_MAX` days into int64
micros and wraps, so e.g. `d < TIMESTAMP 'infinity'` returns the
+infinity row on unpinned / DuckDB-format scans today. The fold gives
the right answer where it engages.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Handle breaking changes in 26.12 to fix nightly job.
…sion exploration (sirius-db#1726)

## Description
This PR brings in the changes from Felipe's late-mat branch sirius-db#1391 to fix
and improve simpatico's explore function.
It includes excluding JIT costs from the evaluation, and outputting all
pareto-optimal points so we can change the throughput floor without
re-measuring.

## Checklist
- [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
- [x] Cover changes with new or existing tests
- [x] Document configuration changes in code and summarize in the
description above
- [x] Update human and agent documentation (README.md, docs/, skills,
CLAUDE.md)

## References

Co-authored-by: Felipe Aramburu <faramburu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…up metadata walk for pin-served scans (sirius-db#1548)

## Description

Two independent frontend fixes, one per commit, that cut redundant
plan-construction work out
of the query path.

### 1. Build the Sirius physical plan once per query

The transparent path built the Sirius GPU physical plan **twice per
query**: once in
`OnFinalizePrepare` purely to validate GPU support (the resulting plan
was discarded), and
again in `PhysicalSiriusExecution::GetDataInternal` for the actual run.
The validated plan is
now handed to `PhysicalSiriusExecution` and consumed by the first
`GetData` (one-shot — the
move empties the slot). Re-executions of the same prepared operator
rebuild from the logical
plan template exactly as before, so data-drift and rebind semantics are
unchanged.

### 2. Defer the row-group metadata walk for pin-served scans

Every plan build constructs one duckdb-native ingestible per `seq_scan`,
whose constructor ran
`prepare_duckdb_native_walk`: `GetPartitionStats` plus a per-row-group
statistics pass per
prunable filter column and per projected varchar column. For a
**pin-served scan its output is
never consumed** — splits come from the pinned entry's chunk plan and
insert-delta bundles, not
from the walk.

`sirius_plan_get` now marks a scan whose MVCC cache-or-CPU guards all
passed with
`mvcc_pin_serves_scan`; the plan generator forwards it as
`defer_metadata_walk`; the ingestible
skips the walk at construction; and the new
`gpu_ingestible::ensure_metadata_prepared()` hook
runs it lazily from the scan manager, on the query thread, for exactly
the scans that did *not*
match a pinned entry at execution time. A scan that actually reads disk
always gets its walk.

The two compound: fix 1 halves the number of plan builds, fix 2 removes
the walk from the
remaining build on pinned workloads.

### Refusal semantics preserved

- The projected-type gate stays **eager in both modes** (factored out as
`unsupported_projected_type_reason`), so unsupported types still refuse
at plan time, where
  the throw is a clean CPU fallback.
- `sirius_plan_get`'s table-level overflow-string probe strictly
dominates the walk's
per-row-group overflow check. For deferred scans the per-row-group
refusal moves to the lazy
walk (a runtime scan error → CPU fallback); for unpinned scans it stays
at construction.
- A lazy-walk failure throws out of `prepare_for_query` and surfaces as
a runtime scan error,
  matching the walk's other runtime refusals.

No new configuration knobs; behavior is unconditional.

## Measurements

**Re-measured 2026-09-16 after rebasing onto `dev` @ `05148c78`**
(RAPIDS 26.08), same GB300 box,
same duckdb-native SF1000 database, unpinned, q1–q19, best-of-3 per
query, arms alternated, each
arm built in its own worktree. Only runs where the process was in the
box's fast mode are
comparable (see the caveat below):

| | dev @ 05148c7 | this branch | |
|---|---|---|---|
| suite total, run A | 36.362 s | 35.907 s | |
| suite total, run B | 36.696 s | 35.922 s | |
| suite total, min | 36.362 s | 35.907 s | **−1.25%** |
| queries equal or faster | | | 19/19 |

Largest movers: q12 −4.2%, q6 −3.6%, q5 −2.2%, q3 −1.9%. No query slower
outside its own run
spread. This matches the −1.05…−1.23% measured on the previous base — as
expected, fix 1 is the
whole effect on an unpinned workload.

**Caveat — bimodal box, not bimodal code.** During this session the
machine put roughly a third of
launched Sirius processes into a persistent "slow mode": cold first
iterations identical, every
warm iteration ~2.6× slower, for the process's whole lifetime, on
*either* binary. Twelve
alternating single-process probes (q6+q12, 3 iterations) split dev 2
fast / 4 slow and this branch
4 fast / 2 slow, with the two binaries equal inside each mode (q6
1.66–1.71 s vs 1.66–1.71 s fast;
4.4–4.5 s vs 4.4–4.5 s slow). Slow-mode runs are excluded from the table
above. The mechanism is not
diagnosed here (best guess: host-pool placement vs page cache); it
predates and is independent of
this PR.

**Also on both arms:** the pre-existing `SIGSEGV` in
`cuMemcpyBatchAsync_v2` under
`decode_duckdb_native_split` now also fires intermittently on q2 (and
q9), not only q20/q3. Dev
bug, reproduces on the base binary.

<details><summary>Original measurements against dev @ 9624f06</summary>



Measured on a GB300 workstation against a duckdb-native TPC-H SF1000
database, comparing this
branch against its base (`dev` @ `9624f063`), each arm built in its own
worktree.

**Unpinned — the case where the walk cannot be skipped and must run
anyway.** Best-of-3 per
query, 4 complete runs on the base and 3 on this branch, arms
alternated:

| | base | this branch | |
|---|---|---|---|
| suite total, min of runs | 37.784 s | 37.389 s | **−1.05%** |
| suite total, median of runs | 37.882 s | 37.416 s | **−1.23%** |
| per-run spread | 37.78–38.15 s | 37.39–37.46 s | non-overlapping |
| queries faster | | | 15/19 |

The two arms' spreads do not overlap, so this is not run-to-run noise.
Largest movers: q6
−4.1%, q12 −3.4%, q19 −2.3%, q4 −2.2%. Two queries moved the other way
(q14 +3.2%, q13 +2.5%)
but both sit inside their own per-run spread.

This is fix 1 acting alone: with nothing pinned, `defer_metadata_walk`
is never set and the
walk runs eagerly on both arms — but the base runs it *twice per scan*
and this branch once.
Counting the per-operator plan-build log marker on a single-scan query
confirms the mechanism:
**6 emissions on the base, 3 on this branch** (3 logical operators × 2
builds vs × 1).

**Pinned.** Verified by log markers that the feature engages as designed
— one
`deferring metadata walk` and one `reusing finalize-validated Sirius
plan` per query, with the
walk skipped entirely. The magnitude of the pinned-workload win is
**not** re-measured here;
the numbers previously quoted on this PR were taken against a superseded
base and have been
removed rather than carried forward.

**Not measured:** the deferred-walk-then-pin-missing path (the walk runs
from
`prepare_for_query` instead of at plan time). Reaching it requires a pin
to disappear between
plan and execute, which I could not force deterministically at scale. It
runs the same function
on the same thread, just later.

**Unrelated pre-existing failure:** q20 segfaults on the unpinned
duckdb-native path on *both*
arms (4/4 runs), and q3 does so intermittently — `SIGSEGV` in
`cuMemcpyBatchAsync_v2` under
`decode_duckdb_native_split`. It reproduces identically on `dev`, so it
is not from this PR;
it is why the table above covers q1–q19.

</details>

## Checklist
- [x] Read CONTRIBUTING.md
- [x] Cover changes with new or existing tests
- [x] Document configuration changes in code and summarize in the
description above
- [ ] Update human and agent documentation (README.md, docs/, skills,
CLAUDE.md) — no
      user-facing behavior or configuration changes to document

## Tests

- New `test/cpp/scan/test_duckdb_native_deferred_walk.cpp`
(`[duckdb_native_deferred_walk]`,
  CPU-side, file-backed duckdb): deferred construction skips the walk;
`ensure_metadata_prepared()` produces a plan identical to the eager one
(row groups, starts,
counts, filter-stat pruning); unsupported types still refuse at
construction in both modes;
overflow-string refusal placement per mode; a failed walk is retried
rather than latched;
  deferred split claims match eager claims.
- **AC-12 lifecycle-test adaptation**
(`test/cpp/integration/test_query_lifecycle_slot.cpp`):
the runtime-window parser required a "Creating sirius physical plan"
line inside every
execution window. With plan reuse the first execution's plan is created
at finalize, so the
parser now also accepts the plan-reuse line as plan provenance. The
acceptance criterion
itself is unchanged — the id-zero check and the cross-execution
operator-id-sequence-equality
check still verify that operator ids restart for every repeated
execution.
- Local: `[duckdb_native_deferred_walk]` + `[duckdb_native_walker]` 10
cases / 101 assertions
  pass; full default suite green (2828 cases, 32.7M assertions).

## References

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…#1767)

This is the first step towards building Sirius as standalone library.

Closes sirius-db#1732.
…capture (sirius-db#1804)

## Description

Sirius emitted every NVTX range into the default domain, so profiles
interleaved Sirius with `libcudf`/`CCCL` and couldn't be filtered to
Sirius alone.

**NVTX `sirius` domain**

- All 80 annotation sites move into a named `sirius` domain.
- The pipeline range spans task boundaries and can't be scoped, so it
switches to `nvtxDomainRangeStartEx`/`End` — semantics unchanged.
- `src/legacy/` is deliberately left on the global domain.

**The domain tag type is declared twice on purpose**

- simpatico also builds standalone, without `src/include` on its include
path.
- `parquet_benchmark` gets `src/include` but not simpatico's include
root.
- So neither header can include the other — collapsing them fails with
`fatal error: codegen/util/nvtx.hpp: No such file or directory`.
- NVTX keys domains by name, so both tag types resolve to one `sirius`
domain.
- The name comes from a single `SIRIUS_NVTX_DOMAIN_NAME` definition set
from `${PROJECT_NAME}`, so it can't drift.

**quent bump**

- `381e1be` → `b21933d0`, for
[quent#736](rapidsai/quent#736), which groups
explicitly named NVTX domains by name at the presentation boundary.

**Quent context ordering**

- cuCascade's ranges gained a `libcucascade` domain in
[cuCascade#196](NVIDIA/cuCascade#196), already
on `dev` — but that domain never reached Quent.
- `nvtx_injection::dispatch` drops events while its `OnceLock` hook is
unset, and that hook is installed by `quent::create_context`, which ran
inside `telemetry_context`'s constructor — i.e. *after* the memory
manager was built. cuCascade's pinned-pool setup, including the
`nvtxDomainCreateA` that names the domain, landed in that dead window.
- The handle stayed valid, so later cuCascade ranges were captured but
unnamed; the UI showed them under `domain0x1`.
- `telemetry_context::create` now takes an already-built
`rust::Box<quent::Context>`, and `SiriusContext::initialize` builds it
first via the new `make_quent_context()`. This keeps the memory manager
and per-GPU device ids the context declaration needs, so no telemetry is
lost.

**Verification**

- TPC-H SF100 (Q1/Q3/Q5) under nsys: one `sirius` domain alongside
`libcucascade`, `libcudf` and `CCCL`; zero ranges left in the default
domain.
- Under Quent: all four domains now appear by name — `libcucascade`
showed as `domain0x1` before the ordering fix.
- Builds clean after rebase onto `dev` (RAPIDS nightly 26.12):
`sirius_extension`, `sirius_unittest`, `parquet_benchmark`, standalone
`simpatico_codegen`.
- `pre-commit run -a` green.

## Checklist
- [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
- [ ] Cover changes with new or existing tests — NVTX domain membership
and Quent capture timing are observable only in a profiler, so both are
verified via the nsys/Quent runs above; the two
`telemetry_context::create` call sites in tests are updated for the new
signature
- [x] Document configuration changes in code and summarize in the
description above
- [ ] Update human and agent documentation — `quent-telemetry.md`
doesn't enumerate domains, so nothing went stale

## References

Refs [cuCascade#196](NVIDIA/cuCascade#196),
[quent#736](rapidsai/quent#736)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…irius-db#1809)

Some clean up:
- silence some build warnings
- remove some obsolete rapids comparability branches
…s-db#1812)

I was going through the simpatico README and found that (1) some of the
commands were outdated and (2) I hit an RMM version bump-related error
(`rmm::mr::set_current_device_resource_ref()` has been removed upstream
in RMM). This PR fixes both these issues.

Here's a concise summary of the build issue and proposed fix from Codex
that I've edited myself

> ### Build issue
> The standalone Simpatico CLI and two tests failed against the upgraded
RAPIDS/RMM stack because they still used the removed
`rmm::mr::set_current_device_resource_ref()` API. The existing guards
also saved the previous resource as a non-owning
`device_async_resource_ref`, which can dangle after the current resource
is replaced.
> 
> ### Fix
> Switch resource installation to `cudf::set_current_device_resource()`,
retain its returned owning `cuda::mr::any_resource`, and move that
resource back when restoring the previous allocator. This follows the
ownership handling introduced in
sirius-db#877, the non-deprecated cuDF
API migration in sirius-db#1098, and the
existing RAII pattern from
sirius-db#1205.
> The change here update the Simpatico CLI, operator-sweep test, and
compression round-trip test. The build and CLI round-trip verification
now pass.

cc @felipeblazing

---------

Signed-off-by: James Bourbeau <jbourbeau@nvidia.com>
## Description

CTrack was fetched and linked unconditionally, which made a
profiling-only dependency part of every build. This adds
`BUILD_WITH_CTRACK`, defaulting to `OFF`, and gates the CTrack fetch,
link dependency, and compile definition behind that option so normal
builds do not require CTrack.

## Validation

- `pixi run pre-commit run -a` passes at this commit.
- The final stacked tree passes the complete release build.

## Checklist

- [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements.
- [x] Preserve the existing CTrack path when `BUILD_WITH_CTRACK=ON`.
- [x] Document the new build option next to its CMake definition.

## Stack

Layer 2 of 6. Based on `stacked/frugal-native-01-rapids`; followed by
sirius-db#1776.

---------

Co-authored-by: Matthijs Brobbel <m1brobbel@gmail.com>
## Description

Replace the two watchdog runners' `fork()` + environment mutation +
`exec()` sequences with `posix_spawn()` using an environment prepared
entirely in the parent process.

Add a reusable child-process environment helper that snapshots the
inherited environment, applies child-only overrides and removals, and
leaves the parent unchanged. This generalizes the existing
envp-construction pattern in `test_operator_sweep.cpp` and
`test_fused_operator_sweep.cpp` so both watchdogs share it.

The spawn environments preserve the original pre-exec behavior,
including inheriting `SIRIUS_DISABLE`. After re-exec, the hidden runners
reapply `SIRIUS_CONFIG_FILE` and clear `SIRIUS_DISABLE` because shared
test-environment startup resets those process-wide values before the
selected test runs.

A tree-wide process-launch audit found no other post-fork environment
mutation. Other fork/exec sites either prepare `envp` in the parent or
perform only descriptor and exec operations in the child, so they are
intentionally outside this change.

## Validation

Validated on Ubuntu 24.04 with an NVIDIA GeForce RTX 5060 Ti (compute
capability 12.0), CUDA 13.2, and `io_uring` enabled:

- `CUDAARCHS=120 CMAKE_BUILD_PARALLEL_LEVEL=8 pixi run make`
- `pixi run pre-commit run -a`
- Child-process environment regression: 7 assertions in 1 test case
- Query-lifecycle watchdog: 78 assertions in 11 test cases
- Hive-partition CUDA watchdog: 28 assertions in 1 test case
- `pixi run make s3-test`: 31,433 assertions in 92 test cases using
testcontainers-managed MinIO over HTTP and TLS
- Native macOS C++17 spawn/environment smoke test passed

## Checklist

- [x] Read CONTRIBUTING.md
- [x] Cover changes with new or existing tests
- [x] Document configuration changes in code and summarize in the
description above (N/A: no configuration changes)
- [x] Update human and agent documentation (N/A: internal test
infrastructure only)

## References

Closes sirius-db#1440

---------

Co-authored-by: Aaron Wu <Aaron.Wu@dell.com>
Reduce dependency on DuckDB-owned standard-library-like types.

- Replace DuckDB vectors and smart pointers for Sirius
pipeline/meta-pipeline ownership with standard-library equivalents.
- Propagate the standard types through pipeline conversion, query
indexing, task creation, telemetry, operator ports, and focused test
helpers.
- Remove direct DuckDB common-pointer/container includes where the
pipeline headers no longer require them.
## Description
There used to be separate compilation modes in Simpatico for decode vs
encode, 1 of them using c++17 and the other c++20. There isn't really a
reason to do so, and it turns out c++20 is also faster.

## Checklist
- [ ] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
- [ ] Cover changes with new or existing tests
- [ ] Document configuration changes in code and summarize in the
description above
- [ ] Update human and agent documentation (README.md, docs/, skills,
CLAUDE.md)

## References

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
## Description

Query lifecycle notifications were coupled to their producers, making it
difficult for readahead and other services to observe execution without
adding direct dependencies. This introduces a query event publisher and
hook-driven subscribers, stamps publications with process-wide event IDs
and system-clock timestamps, and adds the thread utilities used to
dispatch subscriber work safely.

Tests cover event publication, subscriber behavior, ordering metadata,
and the supporting thread utilities.

## Validation

- `pixi run pre-commit run -a` passes at this commit.
- The final stacked tree passes the complete release build.

## Checklist

- [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements.
- [x] Cover the publisher, subscribers, and thread utilities with tests.
- [x] No user-facing configuration change.

## Stack

Layer 3 of 6. Based on `stacked/frugal-native-02-ctrack`; followed by
sirius-db#1774.
…ade/offload copies (sirius-db#1552)

# spill: fix the partition-spill cliff and monolithic downgrade copies
(T8)

**Branch:** `pr/spill-cliff-chunked-offload` → base: `dev` (`3a0ef61f`).
No stacking dependency.

## Problem

The q9-class RF1 spill cliff: +0.1 % delta rows tipped two extra ~4.5 GB
partition round-trips
through host (D2H 57→75 GB / H2D 75→93 GB, post-RF1 +117 ms), delivered
as single 18–28 GB
blocking copy submissions on an un-instrumented thread. Two independent
causes:

1. The builtin cucascade GPU→HOST fast converter submits an entire
batch's D2H copies as one
   monolithic batched call plus a blocking sync.
2. Monitor-issued downgrades force-flush the whole trigger→stop band (≥
12 GB per marginal
crossing on the SF1000 GB300 config) and spill whole partitions in
policy order, so a
   marginal working-set increase becomes a multi-GB spill.

## Change

1. **Chunked, pipelined GPU→HOST spill copies**
(`src/data/spill_chunked_converters.cpp`,
`chunked_spill_copy.hpp`): `SiriusContext::initialize` replaces the
builtin converter with a
Sirius-side transcription that flushes D2H submissions every
`~copy_chunk_bytes` (default
1 GiB; yaml `executor.downgrade.copy_chunk_bytes`, `0` keeps the
builtin) while the column
tree is still being walked — chunk k's DMA overlaps chunk k+1's prep.
Produces a
byte-identical `host_data_representation` (same self-describing
column_metadata layout); the
HOST→GPU restore direction deliberately stays on the unchanged builtin
converter. Adds NVTX
ranges to the downgrade worker's conversions (no timing instrumentation
is compiled into the
   production path; measurement is via the profiler).
2. **Overflow-proportional spill sizing** (`downgrade_executor.cpp`, new
`spill_policy.hpp`):
- Monitor requests (yaml
`executor.downgrade.overflow_proportional_spill`, default true) stop
as soon as live pressure drops back below the *trigger* threshold
instead of flushing down
     to the stop threshold.
- Byte-targeted requests carry `target_bytes`: dispatch stops once
planned (freed +
in-flight) bytes cover the target, bounding overshoot to < 1 batch at
any pool width, and
the final per-repo pick is best-fit (smallest candidate covering the
remaining deficit)
     instead of the next whole partition in policy order.
- The live predicate is evaluated before every dispatch, so an
already-satisfied request
     spills nothing instead of at least one batch.

Reservation-driven `request_downgrade` keeps its
dispatch-until-satisfied envelope
(livelock-sensitive at 251/256 GB); it only gains the earlier predicate
checks. This decouples
spilled bytes from both the trigger→stop band and whole-partition
granularity **without**
retuning `hash_partition_bytes` (a measured trade, not a win).

## Measured

SF1000, GB300: **q9 post-RF1 spill cliff eliminated (+117 ms → −71
ms)**; spilled bytes become
proportional to the actual overflow instead of the trigger→stop band.

## Testing

- New unit tests: `test_chunked_spill_copy.cpp` (chunked converter
produces byte-identical host
representations across chunk-size sweeps), `test_spill_policy.cpp`
(best-fit picks), extended
`test_downgrade_executor.cpp` (predicate-stops, target_bytes bounding).
- Full `sirius_unittest` suite green on GB300: 2842 test cases,
32,778,745 assertions.
- Full pre-commit hook set passes on the changed files.

## Related follow-up (not in this PR)

The stack lineage carries `dcc4c50a` "fix(downgrade): wake the scheduler
matcher after returning
extracted tasks" — a pre-existing whole-process deadlock (task returned
via RAII emits no
`task_available`; deterministic at SF1000 host-pinned q18, 7/7 hangs).
It applies to `dev`
independently of this PR and should be its own small PR.


🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Bobbi Winema Yogatama <34973829+bwyogatama@users.noreply.github.com>
…g, clause-2.8 evolution, scratch rollback, nsys mapping) (sirius-db#1554)

# bench: official TPC-H power/throughput/QphH harness (RF1/RF2,
per-query telemetry + nsys)

Tooling-only PR — no engine code. Branch:
`pr/tpch-power-throughput-harness`, rebased onto current `dev`
(2026-09-16).

## Description

Extends `test/tpch_performance/tpch_power_throughput.py` and the
`bench/sf1000-repro` kit into a
harness for **official TPC-H power/throughput scoring** on a file-backed
`.duckdb` dataset, with
per-query telemetry and per-query nsys profiling for analysis runs. This
harness produced the
SF1000 QphH campaign on GB300 that took **QphH@Size from 5.47M to
7.52M** (reference run in the
commit message: Power@Size 7,165,616 / Throughput@Size 4,951,261 / QphH
5,956,411 on the PR sirius-db#1409
stack, GPU-vs-CPU validation PASS after both refresh functions).

### Scoring (spec-conformant)

- `Power@Size = 3600·SF / geomean(22 stream-0 query times + T_RF1 +
T_RF2)` — the refresh
functions are timed as spec elements of the power geomean, not warmup.
Query times below
`slowest/1000` are floored per **clause 5.4.1.4** (the 1000:1 spread
cap).
- `Throughput@Size = N·22·3600 / measurement_interval · SF` — the
interval is wall-clock from
first stream start to last stream end (**clause 5.3.6.1**), with N
concurrent query streams
  (spec permutations) plus one refresh stream running N RF1/RF2 pairs.
- `QphH@Size = sqrt(Power · Throughput)`.
- `--vary-predicates` runs each stream's own qgen substitution
parameters, as an official run
  requires; the default fixed parameters keep validation meaningful.
- `--update-set-offset` implements **clause 2.8 database evolution**:
one evolving database is
reused across runs (`--scratch-db <shared> --keep-scratch-db`) — the
power run consumes update
set N+1, throughput N+2..N+streams+1 — instead of restoring a 440 GB
pristine copy per
benchmark run. Validation requires offset 0 (`--update-set-offset`
rejects `--validation`);
the scratch size check accepts the evolved size and reports drift
instead of failing.

### Shared-GPU probe gate (`run-power.sh`)

The config reserves ~0.95 of the device at `LOAD`. On a shared box that
fails not only while
another workload runs but **for minutes after it exits** — the driver
reclaims a dead process's
async-pool backing lazily, during which `nvidia-smi` already reports the
memory free but large
allocations still OOM. Gating on `nvidia-smi` therefore passes and the
run still dies; the only
reliable gate is the failing operation itself. `run-power.sh` retries a
throwaway extension-LOAD
probe (`PROBE_TRIES`, default 20 × 60 s) before committing to the run,
on the original
(non-telemetry) config so a Quent capture never sees a second same-named
engine.

### Rollback-scratch: WAL-discard restore (`--rollback-scratch`,
`ROLLBACK=1`)

Restores the scratch DB after a run **without re-copying** (a 15-minute,
440 GB operation at
SF1000). Recipe — three thresholds keep every refresh mutation confined
to the WAL:
`SET checkpoint_threshold='1TB'`, `SET wal_autocheckpoint='1TB'`, and
`SET auto_checkpoint_skip_wal_threshold=2^40` (so bulk writes do not
checkpoint-on-commit). The
run then ends via `os._exit` because DuckDB force-checkpoints the WAL
into the base file on any
clean shutdown and offers no opt-out; deleting `<scratch>.wal`
afterwards (run-power.sh's
`ROLLBACK=1` does it) restores content-pristine state, so every run
reuses one copy at offset 0.

Qualified at small scale: 3 mutate/discard cycles returned exact
pristine content each time.
**Caveat (orphan drift):** the base file keeps small orphaned-block
growth from DuckDB's
optimistic bulk writes — never read, never reclaimed — so the file size
drifts up across cycles;
the size check reports the drift, and the copy should be refreshed when
it grows large.
Incompatible with `NSYS=1`.

### Per-query nsys (`--nsys-per-query`, `NSYS=1`)

Each sequential power-pass query is bracketed in its own cudaProfilerApi
capture range
(`--capture-range=cudaProfilerApi --capture-range-end=repeat::sync`);
the throughput phase gets
one whole-interval range (cudaProfilerStart is process-global —
concurrent streams must not use
it). Two gotchas are baked in:

- **`repeat::sync`, and never trust report indices**: nsys silently
merges adjacent capture
ranges when profiler_stop/start pairs arrive faster than it finalizes a
range, and can drop the
last range at exit — the surviving `range.<N>.nsys-rep` files are
numbered compactly, so
index-based assignment shifts every pointer after the first merge
(observed: 72 files for 89
planned ranges). `prep_analysis_bundles.py` therefore joins reports to
manifest entries by
**absolute-timestamp overlap** (report capture window from its sqlite
export, manifest window
from the runner-recorded `start/stop_epoch_ns`, falling back to the
quent query's Init..Exit
window), and writes a full accounting to `nsys_range_map.json` — mapped
/ ambiguous (merged) /
dropped / no_window / conflict; orphan reports and partial overlaps
flagged, nothing silently
  reassigned. A 0-byte sqlite export is treated as failed, not cached.
- **env-clear for the nsys frontend**: the frontend's bundled
`libssl.so.3` predates the pixi
libcurl's `OPENSSL_3.2.0` requirement and dies at startup, so nsys is
launched with
`LD_PRELOAD`/`LD_LIBRARY_PATH` cleared and both restored only for the
profiled python.

nsys runs are analysis runs; their metrics are never quoted as scores.

### Validation worker design

Validation cannot share a process with the pinned run: Sirius's host
pool is a growing pool
allocator (unpinning returns blocks to the pool, not the OS) and DuckDB
sizes its default
`memory_limit` from total system RAM with no knowledge of what Sirius
holds — at SF1000 a
CPU-side q9 allocates into a machine already ~320 GB spoken for and gets
OOM-killed. The GPU rows
are pickled per phase, and a **fresh child process** that never loads
the extension replays
RF1/RF2 on its own copy of the pristine input to reproduce the
post-refresh states, then diffs
row-by-row (sorted multisets, absolute tolerance). Related:
`--duckdb-memory-limit` (default
32 GB) caps the benchmark connection itself — the unlimited default
OOM-killed the pinned process
during pin materialization at SF1000.

### Telemetry and analysis prep

- Every query is labeled (`CALL sirius_set_query_label`, zero-padded
`q01`..`q22`) and each phase
bucketed into its own Quent query group via `CALL
sirius_set_session_label` (`warmup`,
`power_clean`, `power`, `power_postrf2`, `tput_s1..sN`); both calls sit
outside the timed
  window. `QUENT=1` derives a telemetry-enabled config per run.
- `prep_analysis_bundles.py <run_dir> <nsys_dir> <quent_dir>` builds
per-query analysis bundle
JSONs (nsys report pointers with `nsys_status`/`nsys_shared_with`,
timings, quent UUIDs) and
  pre-exports every `.nsys-rep` to sqlite in parallel.
- `--pin-layout <json>`: mixed-tier pin layouts with validated column
coverage (a
schema-qualified name creates a second entry over the same table for
per-column tiering);
SF1000 reference layout in `bench/sf1000-repro/pin-layout-sf1000.json`.
- `--warmup-pass` burns one discarded pass so JIT/first-touch costs land
nowhere; crash-safe
cleanup keeps results when unpin fails; the perf-stack env
(`SIRIUS_EXP_*`, `SIRIUS_PRE_SQL`,
`LD_PRELOAD`) is recorded into `run_info.txt`/`metrics.json` for
provenance.
- `generate_tpch_refresh.sh`: dbgen `upath[128]` overflow fix (stage via
`/tmp`).

## Engine dependencies (all in `dev` now)

Everything this harness leans on has landed, and the rebase commit
adapts the harness to it:

- **`sirius_set_session_label`** exists in `dev`; the runner still
catches `CatalogException` and
  degrades to unlabeled groups on an older engine.
- **Split pin entries** (the `main.orders` host-tier entry in the SF1000
layout) use the
column-aware `find_pinned_entry_for_duckdb_table(requested_ids)` lookup
that is in `dev`. A
scored run must still have zero `Transparent execution fallback` lines
in its `log_dir`.
- **`SIRIUS_EXP_*` gates** (sirius-db#1524): `run-power.sh` exports the same
names and defaults as `run.sh`
(`SIRIUS_EXP_FUSED_SCAN_FILTER`, `SIRIUS_EXP_LATE_MAT`,
`SIRIUS_EXP_LATE_MAT_PIN_UNIQUE_COLS`).
Late-mat is inert on duckdb pins; fused scan-filter engages on the
compressed GPU pins.
- **Config** (sirius-db#1499): the two diagnostic YAMLs no longer set
`thread_name_prefix`, which the loader
  now rejects.
- **Patched libcudf** is optional (`CUDF_SO`), matching `run.sh` after
sirius-db#1524; unset uses the pixi
  libcudf.
- `SIRIUS_POISON_FREES` / `SIRIUS_QUARANTINE_FREES` (used by
`verify-poison-concurrent.sh`) are the
one remaining separate piece; without them that gate degrades to a
concurrent stress run with
  coredump capture.

## Reconciliation with sirius-db#1470

PR sirius-db#1470 (fadvise cache eviction, benchmark totals) landed in
`test/tpch_performance/performance_test.py`; this PR does not touch that
file and keeps dev's
shape. The runner's own `_evict_page_cache` uses the same mechanism
(`posix_fadvise(DONTNEED)`,
no sudo) but targets the `.duckdb` scratch/input files with an
fsync-first dirty-page path — it
complements rather than duplicates `drop_os_cache`.

## Staged refresh (sirius-db#1551)

Staged refresh landed in `dev` while this PR was open and is on by
default in the runner; the
clause-2.8 offset and the rollback path both pass it through
(`rf1/rf2_statements(dir, set,
staged)`), and `run-power.sh` forwards `--no-staged-refresh` for the
legacy path. Staging tables
are created after the rollback thresholds are armed, so they live in the
WAL and are discarded
with it.

## Rollback on every exit path

The rebase commit also closes a gap in `--rollback-scratch`: an
exception mid-run used to reach
interpreter teardown, where the kept-alive connection's destructor
checkpoints the WAL into the
base file and consumes the scratch — on exactly the crashed run whose
scratch needs restoring.
`os._exit` now fires on every exit path once rollback is armed, and
`run-power.sh` reaches its
WAL-discard step even when the runner fails (`|| RC=$?` instead of
tripping `set -e`).

## Checklist

- [x] Read CONTRIBUTING.md
- [x] Cover changes with new or existing tests — `py_compile`/`bash
-n`/`pre-commit` clean;
exercised end-to-end at SF1000 (see measured context above) and at small
scale for the
rollback qualification. No unit tests: benchmark tooling, not engine
code.
- [x] Document configuration changes — knob table in
`bench/sf1000-repro/README.md`, flag docs
      in `test/tpch_performance/CLAUDE.md`.
- [x] Update human and agent documentation (README.md, CLAUDE.md)

## References

Refs sirius-db#1470 (cache-eviction overlap, reconciled above).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Description

Components that finish work on CUDA streams need to observe event
completion without blocking an executor thread. This adds
`cuda_event_completion_poll`, which polls CUDA events and dispatches
their completion callbacks through the execution infrastructure.

The accompanying tests cover registration, polling, completion dispatch,
shutdown, and error paths.

## Validation

- `pixi run pre-commit run -a` passes at this commit.
- The final stacked tree passes the complete release build.

## Checklist

- [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements.
- [x] Cover normal completion and lifecycle edge cases with tests.
- [x] No user-facing configuration change.

## Stack

Layer 4 of 6. Based on `stacked/frugal-native-03-query-events`; followed
by sirius-db#1773.
…#1773)

## Description

The I/O code used prefetch-specific names for objects that now represent
cache residency and completion more generally. This normalizes the I/O
namespace and source layout, moves REST and S3 helpers under `io/rest`,
introduces the shared invocable execution utility, and renames
`prefetching_handle` to `cache_handle` throughout scan, cache, and
engine code.

The rename prepares the interfaces consumed by the dynamic I/O and
frugal-caching layer while keeping behavior unchanged.

## Validation

- `pixi run pre-commit run -a` passes at both commits in this PR.
- The final stacked tree passes the complete release build.

## Checklist

- [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements.
- [x] Update affected tests and documentation with the normalized names.
- [x] No user-facing configuration change.

## Stack

Layer 5 of 6. Based on `stacked/frugal-native-04-cuda-events`; followed
by sirius-db#1771.
…#1593)

This PR simply updates sirius to be able to compile against updates in
NVIDIA/cuCascade#192, so we can incrementally
move to a more bolted down data management strategy.

This is intentionally narrowly scoped to ease the review process, and
the subsequent independent pieces will be proposed separately.

Part 1 of a few to move towards solving
sirius-db#1063 from the root of it.
…sirius-db#1828)

Bumps [clap](https://github.com/clap-rs/clap) from 4.6.6 to 4.6.7.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/clap-rs/clap/releases">clap's
releases</a>.</em></p>
<blockquote>
<h2>v4.6.7</h2>
<h2>[4.6.7] - 2026-09-14</h2>
<h3>Features</h3>
<ul>
<li><em>(derive)</em> Add <code>#[command(defer = &lt;bool&gt;)]</code>
attribute to opt-in to lazy initialisation of subcommands</li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/clap-rs/clap/blob/main/CHANGELOG.md">clap's
changelog</a>.</em></p>
<blockquote>
<h2>[4.6.7] - 2026-09-14</h2>
<h3>Features</h3>
<ul>
<li><em>(derive)</em> Add <code>#[command(defer = &lt;bool&gt;)]</code>
attribute to opt-in to lazy initialisation of subcommands</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/clap-rs/clap/commit/d3e59a9ab214910b9dad02921b7ef42c6400de9b"><code>d3e59a9</code></a>
chore: Release</li>
<li><a
href="https://github.com/clap-rs/clap/commit/d997f878c484e02b2935d32a7b67afe990e91227"><code>d997f87</code></a>
docs: Update changelog</li>
<li><a
href="https://github.com/clap-rs/clap/commit/fb6058cc38cf6f4b07b8db89c850fc80e059e222"><code>fb6058c</code></a>
Merge pull request <a
href="https://redirect.github.com/clap-rs/clap/issues/6409">#6409</a>
from heaths/pwsh-support</li>
<li><a
href="https://github.com/clap-rs/clap/commit/2310870f7a9fe4c39f5bb5620261a06ac9aaa0b1"><code>2310870</code></a>
test(complete): Add tests for completer_for_path</li>
<li><a
href="https://github.com/clap-rs/clap/commit/5967c17393a09f90112cdc5b2c20a1383f0bd8e4"><code>5967c17</code></a>
refactor(complete): Move shell detection to Shells</li>
<li><a
href="https://github.com/clap-rs/clap/commit/594602bb29f374e49da5029df63e5c740c723ad0"><code>594602b</code></a>
fix(complete): Detect pwsh for PowerShell</li>
<li><a
href="https://github.com/clap-rs/clap/commit/3a4f2d031b5110aaaac40f1cb1d6c0b8ff619df8"><code>3a4f2d0</code></a>
Merge pull request <a
href="https://redirect.github.com/clap-rs/clap/issues/6427">#6427</a>
from clap-rs/renovate/shlex-2.x</li>
<li><a
href="https://github.com/clap-rs/clap/commit/67ebaed4ac95a83d5d7cc22d64ee0fc162ae1db3"><code>67ebaed</code></a>
Merge pull request <a
href="https://redirect.github.com/clap-rs/clap/issues/6426">#6426</a>
from clap-rs/renovate/actions-checkout-7.x</li>
<li><a
href="https://github.com/clap-rs/clap/commit/c968b136d0e6f560c08b06a844a20f2798a00096"><code>c968b13</code></a>
chore(deps): Update Rust crate shlex to v2</li>
<li><a
href="https://github.com/clap-rs/clap/commit/8f247cbf58c582a755ee4dbe62a2390d93b9f8eb"><code>8f247cb</code></a>
chore(deps): Update actions/checkout action to v7</li>
<li>Additional commits viewable in <a
href="https://github.com/clap-rs/clap/compare/clap_complete-v4.6.6...clap_complete-v4.6.7">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=clap&package-manager=cargo&previous-version=4.6.6&new-version=4.6.7)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [clap](https://github.com/clap-rs/clap) from 4.6.6 to 4.6.7.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/clap-rs/clap/releases">clap's
releases</a>.</em></p>
<blockquote>
<h2>v4.6.7</h2>
<h2>[4.6.7] - 2026-09-14</h2>
<h3>Features</h3>
<ul>
<li><em>(derive)</em> Add <code>#[command(defer = &lt;bool&gt;)]</code>
attribute to opt-in to lazy initialisation of subcommands</li>
</ul>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/clap-rs/clap/blob/main/CHANGELOG.md">clap's
changelog</a>.</em></p>
<blockquote>
<h2>[4.6.7] - 2026-09-14</h2>
<h3>Features</h3>
<ul>
<li><em>(derive)</em> Add <code>#[command(defer = &lt;bool&gt;)]</code>
attribute to opt-in to lazy initialisation of subcommands</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/clap-rs/clap/commit/d3e59a9ab214910b9dad02921b7ef42c6400de9b"><code>d3e59a9</code></a>
chore: Release</li>
<li><a
href="https://github.com/clap-rs/clap/commit/d997f878c484e02b2935d32a7b67afe990e91227"><code>d997f87</code></a>
docs: Update changelog</li>
<li><a
href="https://github.com/clap-rs/clap/commit/fb6058cc38cf6f4b07b8db89c850fc80e059e222"><code>fb6058c</code></a>
Merge pull request <a
href="https://redirect.github.com/clap-rs/clap/issues/6409">#6409</a>
from heaths/pwsh-support</li>
<li><a
href="https://github.com/clap-rs/clap/commit/2310870f7a9fe4c39f5bb5620261a06ac9aaa0b1"><code>2310870</code></a>
test(complete): Add tests for completer_for_path</li>
<li><a
href="https://github.com/clap-rs/clap/commit/5967c17393a09f90112cdc5b2c20a1383f0bd8e4"><code>5967c17</code></a>
refactor(complete): Move shell detection to Shells</li>
<li><a
href="https://github.com/clap-rs/clap/commit/594602bb29f374e49da5029df63e5c740c723ad0"><code>594602b</code></a>
fix(complete): Detect pwsh for PowerShell</li>
<li><a
href="https://github.com/clap-rs/clap/commit/3a4f2d031b5110aaaac40f1cb1d6c0b8ff619df8"><code>3a4f2d0</code></a>
Merge pull request <a
href="https://redirect.github.com/clap-rs/clap/issues/6427">#6427</a>
from clap-rs/renovate/shlex-2.x</li>
<li><a
href="https://github.com/clap-rs/clap/commit/67ebaed4ac95a83d5d7cc22d64ee0fc162ae1db3"><code>67ebaed</code></a>
Merge pull request <a
href="https://redirect.github.com/clap-rs/clap/issues/6426">#6426</a>
from clap-rs/renovate/actions-checkout-7.x</li>
<li><a
href="https://github.com/clap-rs/clap/commit/c968b136d0e6f560c08b06a844a20f2798a00096"><code>c968b13</code></a>
chore(deps): Update Rust crate shlex to v2</li>
<li><a
href="https://github.com/clap-rs/clap/commit/8f247cbf58c582a755ee4dbe62a2390d93b9f8eb"><code>8f247cb</code></a>
chore(deps): Update actions/checkout action to v7</li>
<li>Additional commits viewable in <a
href="https://github.com/clap-rs/clap/compare/clap_complete-v4.6.6...clap_complete-v4.6.7">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=clap&package-manager=cargo&previous-version=4.6.6&new-version=4.6.7)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
I was going through the Python API section of the README and noticed it
fails to build due to lack of a recent git tag

<details>
<summary>Full traceback:</summary>

```
(base) jbourbeau@rl-b5652-03:~/sirius-python-build$ pixi run -e duckdb-python build-duckdb-python
✨ Pixi task (build-duckdb-python in duckdb-python): CMAKE_ARGS="-DDUCKDB_SOURCE_PATH=$PIXI_PROJECT_ROOT/duckdb" pip install --no-build-isolation ./duckdb-python
Processing ./duckdb-python
  Preparing metadata (pyproject.toml) ... error
  error: subprocess-exited-with-error

  × Preparing metadata (pyproject.toml) did not run successfully.
  │ exit code: 1
  ╰─> [77 lines of output]
      [version_scheme] version object: <ScmVersion 0.0.1.dev1 dist=1 node=gb236c8194ed14c7a7c685e0534dde501cc855b3a dirty=False branch=HEAD>
      [version_scheme] version.tag: 0.0.1.dev1
      [version_scheme] version.distance: 1
      [version_scheme] version.dirty: False
      Traceback (most recent call last):
        File "/home/jbourbeau/sirius-python-build/duckdb-python/duckdb_packaging/setuptools_scm_version.py", line 56, in version_scheme
          return _bump_dev_version(str(version.tag), distance)
        File "/home/jbourbeau/sirius-python-build/duckdb-python/duckdb_packaging/setuptools_scm_version.py", line 73, in _bump_dev_version
          major, minor, patch, post, rc = parse_version(base_version)
                                          ~~~~~~~~~~~~~^^^^^^^^^^^^^^
        File "/home/jbourbeau/sirius-python-build/duckdb-python/duckdb_packaging/_versioning.py", line 33, in parse_version
          raise ValueError(msg)
      ValueError: Invalid version format: 0.0.1.dev1 (expected X.Y.Z, X.Y.Z.rcM or X.Y.Z.postN)

      The above exception was the direct cause of the following exception:

      Traceback (most recent call last):
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/pip/_vendor/pyproject_hooks/_in_process/_in_process.py", line 389, in <module>
          main()
          ~~~~^^
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/pip/_vendor/pyproject_hooks/_in_process/_in_process.py", line 373, in main
          json_out["return_val"] = hook(**hook_input["kwargs"])
                                   ~~~~^^^^^^^^^^^^^^^^^^^^^^^^
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/pip/_vendor/pyproject_hooks/_in_process/_in_process.py", line 175, in prepare_metadata_for_build_wheel
          return hook(metadata_directory, config_settings)
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/build/__init__.py", line 96, in prepare_metadata_for_build_wheel
          return _build_wheel_impl(
                 ~~~~~~~~~~~~~~~~~^
              None, config_settings, metadata_directory, editable=False
              ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
          ).wheel_filename  # actually returns the dist-info directory
          ^
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/build/wheel.py", line 177, in _build_wheel_impl
          return _build_wheel_impl_impl(
              wheel_directory,
          ...<5 lines>...
              pyproject=pyproject,
          )
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/build/wheel.py", line 234, in _build_wheel_impl_impl
          metadata = get_standard_metadata(pyproject, settings)
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/build/metadata.py", line 44, in get_standard_metadata
          new_pyproject_dict["project"] = process_dynamic_metadata(
                                          ~~~~~~~~~~~~~~~~~~~~~~~~^
              new_pyproject_dict["project"], settings.metadata
              ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
          )
          ^
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/builder/_load_provider.py", line 150, in process_dynamic_metadata
          return dict(settings)
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/builder/_load_provider.py", line 114, in __getitem__
          self.project[key] = provider.dynamic_metadata(  # type: ignore[call-arg]
                              ~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
              key, self.settings[key]
              ^^^^^^^^^^^^^^^^^^^^^^^
          )
          ^
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/scikit_build_core/metadata/setuptools_scm.py", line 30, in dynamic_metadata
          version = _get_version(config, force_write_version_files=True)
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/setuptools_scm/_get_version.py", line 29, in _get_version
          return _get_version_core(
              config, force_write_version_files=force_write_version_files
          )
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/vcs_versioning/_get_version_impl.py", line 171, in _get_version
          return _finalize(scm_version, config, force_write=force_write_version_files)
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/vcs_versioning/_get_version_impl.py", line 53, in _finalize
          version_string = _format_version(applied)
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/vcs_versioning/_version_schemes/__init__.py", line 113, in format_version
          main_version = _entrypoints._call_version_scheme(
              version_for_scheme,
              "setuptools_scm.version_scheme",
              version.config.version_scheme,
          )
        File "/home/jbourbeau/sirius-python-build/.pixi/envs/duckdb-python/lib/python3.14/site-packages/vcs_versioning/_entrypoints.py", line 87, in _call_version_scheme
          result = scheme(version)
        File "/home/jbourbeau/sirius-python-build/duckdb-python/duckdb_packaging/setuptools_scm_version.py", line 59, in version_scheme
          raise RuntimeError(msg) from e
      RuntimeError: Failed to bump version: Invalid version format: 0.0.1.dev1 (expected X.Y.Z, X.Y.Z.rcM or X.Y.Z.postN)
      [end of output]

  note: This error originates from a subprocess, and is likely not a problem with pip.
error: metadata-generation-failed

× Encountered error while generating package metadata.
╰─> from file:///home/jbourbeau/sirius-python-build/duckdb-python
```

</details>


This PR removed `--depth=1` in the submodule init command to ensure we
pull the full git history (including tags). I also added a small CI
build step where we ensure the Python API builds successfully (similar
conditions as the rust binding checks).

Signed-off-by: James Bourbeau <jbourbeau@nvidia.com>
Migrate to `cuda::stream_ref`. This fixes our nightly build and also
gets rids of the deprecation warnings.
## Description
At scale factor 3000 on a smaller machine, the new dense join operator
(sirius-db#1606) goes off the rails and
tries to reserve 800GB of device memory (10x the size of the device).
This PR changes the estimation to hopefully track actual usage a bit
closer (please have a look if you think there are situations where the
estimate is now too low!). And it adds a fallback path for when the new
operator will not fit.

Adds config parameter dense_count_join_memory_fraction to set what part
of the memory to use for the dense count join. Default 10%. The other
parameter, setting an explicit byte count for the threshold, still
exists but is now set to 0 (auto) to use the new fraction.

## Checklist
- [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
- [x] Cover changes with new or existing tests
- [x] Document configuration changes in code and summarize in the
description above
- [x] Update human and agent documentation (README.md, docs/, skills,
CLAUDE.md)

## References

---------

Co-authored-by: Joost Hoozemans <1442581+joosthooz@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
We can unify our instructions now that claude also reads them.
This allows us to drop the `protoc` dependency requirement.
buzzcrow and others added 27 commits September 23, 2026 13:27
## Description

Add GPU scan support for the following Parquet virtual columns:

- `filename`
- `file_index`
- `file_row_number`

This PR:

- preserves file identity and file-local row offsets in Parquet split
metadata;
- reads selected row groups as provenance-preserving runs, synthesizes
the requested virtual columns on GPU, and concatenates the results;
- reuses the Parquet batch-layout calculation for both `file_row_number`
and Iceberg positional-delete processing;
- evaluates predicates involving virtual columns after virtual-column
materialization;
- disables row-level cuDF predicate pushdown when virtual columns are
requested, while retaining physical row-group statistics pruning;
- supports virtual-only scans, projection/reordering, duplicate outputs,
Hive partition columns, legacy virtual-column options, and Iceberg
scans;
- bypasses pinned Parquet cache entries for virtual-column queries
because existing pins do not store per-row provenance.

There are no new configuration options.

### file_index compatibility

Sirius keeps `file_index` as the zero-based index in the original bound
file list. For a bound list `[A, B]`, B retains index 1 when a runtime
join filter removes A.

DuckDB can renumber B to 0 when `file_index` is projected and runtime
filtering targets either a Hive partition column or the legacy filename
column exposed by `filename=true`. The standard virtual `filename`
column alone does not enable this pruning. The upstream inconsistency is
tracked in duckdb/duckdb#26044 and is also
reproducible in the standard DuckDB CLI v2.0.0-dev3 (`ed599a4425`).

The runtime-pruning compatibility test covers both triggers. If DuckDB
adopts stable numbering, the test should require CPU/GPU equality and
remove its known-difference expectation.

### Follow-up work

- Use cuDF 26.08.01's existing `enable_prepend_source_index_column()`
and `enable_prepend_row_index_column()` APIs to read ordinary Parquet
splits in one invocation while retaining physical reader filtering and
dynamic-filter AST merging. Map source indices to bound file identities,
adapt the column layout, convert row indices to BIGINT, and update
memory accounting and provenance tests.
- Translate predicates on `filename` and `file_index` into source
selection.
- Adapt Iceberg positional-delete matching to mapped file identities and
original row indices before enabling reader filtering for that path.
- Add explicit `pin_table` support for Parquet virtual columns.
- Improve physical carrier selection and normalization, and determine
when virtual-only scans can safely avoid physical reads.
- Include the filename scalar's device buffer (`path.size()` bytes) in
construction memory accounting.

### Validation

Previously reported local validation (not a final-head CI result):

```text
[scan][parquet][assembly]
6 test cases, 36 assertions passed

[integration][gpu_execution][scan][virtual_columns][acceptance]
8 test cases, 7195 assertions passed

[integration][gpu_execution][hive_partition]
7 test cases, 366 assertions passed

```

The full C++ test suite was also run. Two unrelated REST subprocess
tests initially hit local GPU OOM and both passed on retry.

The 14 previously hidden integration cases are now visible to the
default test run. Final-head CI must pass and exercise them, the
runtime-pruning compatibility case, and the filename sizing cases.

## Checklist

- [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
- [x] Cover changes with new or existing tests
- [x] Document configuration changes in code and summarize in the
description above — no configuration changes
- [x] Update human and agent documentation — no new configuration or
agent-facing interface

## References

Refs sirius-db#492
…#1874)

## Description

Fix the arm64 nightly build failure introduced by sirius-db#1846. The new Parquet
virtual-column helpers declared `rmm::cuda_stream_view` without
including its declaration, causing compilation to fail in
`scan_plan.hpp` and `scan_plan.cpp`.

Use `cuda::stream_ref` in the helpers and their callers, matching the
surrounding scan API. Update the scan output assembly test helper to use
the same type.

  Verification:
- Compiled the affected scan objects and scan output assembly test
source on x86.
  - Commit hooks passed, including clang-format.
- The arm64 nightly build has not been run locally; CI will verify that
target.

  ## Checklist
- [x] Read CONTRIBUTING.md and ensure PR meets the reviewability
checklist
  - [x] Cover changes with existing scan tests
- [x] Document configuration changes in code and summarize them above —
no configuration changes
- [x] Update human and agent documentation — no documentation changes
needed for this build fix

  ## References
  - Follow-up to sirius-db#1846
This adds the `sirius` Rust crate docs to our GitHub pages deploy (in
addition to sirius-db#1860).
…-db#1877)

## Description
Land the prep work sirius-db#1859's rename cutover depends on, so main-triggered
CI
and contributor-facing docs are correct before the freeze starts: CI
workflows now trigger on main alongside dev, CONTRIBUTING.md gains a
local-clone migration guide covering every remote layout in use
(personal
fork, Stacked-PR-maintainer origin, fallback), dry-run verified against
real git behavior, and stray dev-branch references across docs and one
source comment now point at main.

**NOTE:** There will be a follow-up PR once the rename happens on `main`
to remove any CI references for `dev`. These are needed with `main` to
work across the rename

## Checklist
- [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
- [ ] Cover changes with new or existing tests
- [ ] Document configuration changes in code and summarize in the
description above
- [x] Update human and agent documentation (README.md, docs/, skills,
CLAUDE.md)

## References
Refs sirius-db#1859
…ius-db#1878)

## Description
Follow-up to sirius-db#1877. The dev-to-main rename (sirius-db#1859) is complete and `dev`
no longer exists as a branch, so the temporary dual `dev`/`main` CI
triggers added as prep work are no longer needed. `check.yml`,
`test.yml`, `distribution.yml`, and `docs.yml` now trigger on `main`
only, and the two branch-name conditionals in `distribution.yml` and
`docs.yml` drop their `dev` half.

A repo-wide sweep found no other leftover `dev`-as-branch language that
needs cleanup — the remaining mentions in `CONTRIBUTING.md`'s migration
guide and `AGENTS.md`'s pointer sentence are intentional and still
needed until every contributor's local clone/fork/stack has migrated.

**NOTE:** This PR is non-blocking so it can be merged once it has been
reviewed

## Checklist
- [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
- [ ] Cover changes with new or existing tests
- [ ] Document configuration changes in code and summarize in the
description above
- [x] Update human and agent documentation (README.md, docs/, skills,
CLAUDE.md)

## References
Closes sirius-db#1859
…og lowering/plan/execute timings (sirius-db#1841)

## Why

Three gaps on the FFI path (`sirius_ffi.cpp`, the embedded engine used
by `rust/crates/sirius-sys` and the `experimental/` backends), all found
while profiling TPC-H SF100 through it:

1. **Substrait lowering re-parsed parquet footers on every bind.** The
Substrait consumer builds the plan through the Relation API, which
re-binds the whole subtree at every level, so a parquet read is bound
once per operator above it. With the metadata cache off each bind
re-parses the footer and re-derives the column statistics. On a 25 GB,
~5k-row-group SF100 `lineitem`, lowering took 1.3 s for Q6 and **10 s
for Q21 — as long as or longer than the GPU execution (1.3 s / 4.6 s)**.
2. **No log output.** The embedded DuckDB never loads the Sirius
extension, so nothing on this path read `SIRIUS_LOG_{BACKEND,DIR,LEVEL}`
the way `SiriusContextExtensionCallback` does on the transparent path;
the engine ran with the noop sink and the host process could not get a
Sirius log.
3. **No per-phase timing.** The host only sees the total
`execute_substrait` time, and the telemetry query window covers
execution alone, so time spent in lowering/planning (like the 10 s
above) was invisible.

## What

- `bring_up()`: set `parquet_metadata_cache = true` on the embedded
DuckDB **at the database level** (`DBConfig::SetOptionByName`). It has
to be database-level: the consumer's binder does not see session-level
`SET` variables. This DuckDB instance is private to the engine and the
cache validates file mtimes, so there is no staleness hazard for the
host. Lowering drops to 55–220 ms on the cases above.
- `install_log_sink_from_env()`: at bring-up, if any of
`SIRIUS_LOG_BACKEND` / `SIRIUS_LOG_DIR` / `SIRIUS_LOG_LEVEL` is set,
copy them into `duckdb::Config::LOG_*` and call
`install_configured_log_sink(nullptr)` (best-effort, an unknown backend
is ignored rather than failing bring-up). With none of the variables set
the sink is left untouched, i.e. behaviour is unchanged by default.
- `execute_substrait()`: time the lowering / physical-planning /
execution phases and emit one `SIRIUS_LOG_INFO` line per query.

+52/−1, `sirius_ffi.hpp` unchanged, no API change. Applies to every FFI
consumer (StarRocks backend, `sirius-sys`), all in the "faster and
observable" direction.

## Verification

- TPC-H SF1 22/22 through the FFI path unchanged vs the DuckDB baseline.
- SF100 on an L40S (`g6e.8xlarge`): per-query engine time drops by the
lowering delta (Q21 10 s → ~0.2 s lowering); the new log line is what
exposed the split.
- Existing `test/cpp/exec/test_sirius_ffi_fragment.cpp` covers
bring-up/execute on this path; no new test since the change is
configuration + logging.

## Relationship to other PRs

Independent of the join fix (sirius-db#1840). sirius-db#1842's benchmark numbers were
taken with this change in place.

---------

Co-authored-by: morningman <moringman@apache.org>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ize inequality sides (sirius-db#1840)

## Why

Two failures that only the Substrait/FFI path (`sirius_ffi.cpp` →
`sirius_physical_plan_generator`) reaches. The transparent
DuckDB-extension path plans the same SQL differently — DELIM_JOIN keeps
the inequality in a filter, and the optimizer puts the `NOT IN` mark
join on top — so it never hit them. Found by running TPC-H through the
FFI path (the plans an Apache Doris FE produces, see sirius-db#1842) on a T4.

## What

**`src/op/sirius_physical_hash_join.cpp` — MARK join polled before
sizing (TPC-H Q16, process abort).**
A MARK join whose probe pipeline is the source of an inner join's probe
side gets polled by the task creator (`WAITING_FOR_INPUT_DATA` walk)
before either of its input partitions negotiated `BUILD_PROBE`.
`get_next_task_hint` then takes the STANDARD/MIXED branch and
`refresh_cross_schedule` throws its MARK invariant from the task-creator
thread, which terminates the process. A MARK join always ends up in
`BUILD_PROBE`, so when polled in that window there is nothing to
schedule yet: return `WAITING_FOR_INPUT_DATA` on the build producer
instead, whose tasks feed the build partition and size the join.

**`src/planner/sirius_plan_comparison_join.cpp` — DECIMAL on an
inequality join side (TPC-H Q17/Q20, runtime error).**
`materialize_expression_join_keys` skipped inequality conditions
entirely, leaving both sides to the mixed join's inline cuDF AST
predicate. cuDF AST can neither cast to nor compute with DECIMAL, so
`cast(l_quantity as decimal(38,5)) < 0.2 * avg(l_quantity)` failed with
*"failed to translate mixed join inequality conditions to cuDF AST
predicate"*. Inequality sides now go through the same materialization as
routed null-safe keys: anything beyond a column reference or an
AST-supported cast is projected into a column below the join, leaving a
plain column comparison for the predicate. Equality keys are unchanged.

Net: +23/−7, no header or public-API change, no behaviour change for
plans that already worked (the MARK early-return only fires in the
previously-throwing window; materialization only kicks in for
expressions the AST could not evaluate).

## Verification

- TPC-H SF1, all 22 queries, through the FFI path **and** through the
transparent path: results match the DuckDB baseline before/after (the
transparent path exercises the unchanged code paths for regression).
- Q16 no longer aborts; Q17/Q20 no longer fail at runtime on the FFI
path.
- No new unit test in this PR: the failing shapes only come out of the
Substrait consumer. Happy to add a
`test/cpp/exec/test_sirius_ffi_fragment.cpp`-style case if a reviewer
wants one — say so and I'll push it.

## Relationship to other PRs

Independent of the other two PRs from this branch split. sirius-db#1842
(`experimental/doris`) needs this one at runtime for Q16/Q17/Q20 but
does not build against it.

---------

Co-authored-by: morningman <moringman@apache.org>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…irius-db#1818) (sirius-db#1843)

## Description

Refactors the existing dynamic-filter implementation without adding
multi-batch build support.

- Moves filter publication out of the hash join into a session that
handles completion, cancellation, and cleanup.
- Tracks which producers can still publish filters, so consumers can
tell when no more filters will arrive.
- Gives each filtering pass a consistent snapshot for choosing filters,
applying them, and measuring their usefulness.
- Keeps filter storage alive while GPU work uses it and adds tests for
these lifetimes.

Build data is still delivered before filters are constructed and copied
to other GPUs. There are no cuCascade changes. No performance comparison
is included.

## Checklist
- [x] Read CONTRIBUTING.md and ensure PR meets "reviewability" checklist
- [x] Cover changes with new or existing tests
- [x] Document configuration changes in code and summarize in the
description above
- [x] Update human and agent documentation (README.md, docs/, skills,
CLAUDE.md)

## References
Addresses sirius-db#1820 of issue tracker sirius-db#1818
Next towards sirius-db#1734. This one introduces a shared object target.

Closes sirius-db#1733.
We don't really need to keep and track `duckdb-python` in our repo. The
extension we build should just load in a compatible DuckDB's Python
package.
Next step in sirius-db#1734. This splits out the DuckDB dep into it's own target.
This later allows us to fully own the build i.e. no DuckDB extension
build infra in the main project.
This fixes a few more deprecation warnings (that show up in the nightly
build).
Follow-up of sirius-db#1870. Replacing `cmake/sirius-sources.cmake` with
`CMakeLists.txt` files per subtree, included via `add_subdirectory`.
## Description

Large and remote parquet scans need to avoid eagerly reading and
decoding data that a query may never consume. This adds frugal caching
and dynamic I/O so physical requests can be split, aligned, coalesced,
staged, and served incrementally according to each backend's
capabilities.

The layer includes:

- partial-read cache fills with reader/evictor arbitration and explicit
lifecycle handling;
- reactor-owned request splitting and bulk prepared-slice APIs for REST,
io_uring, and kvikIO;
- query-event-driven readahead, scheduling, and scan-manager
integration;
- parquet materialization and device-copy support;
- `reset_sirius_cache()` and updated cache/backend configuration;
- `pin_table(..., tier='parquet')` for retaining undecoded parquet
ranges in the I/O cache;
- updated documentation, performance scripts, and I/O benchmarks.

Tests cover cache state transitions, request planning, backend behavior,
readahead lifecycle, configuration, reset behavior, and parquet scan
sizing.

## Validation

- `pixi run pre-commit run -a`
- `pixi run cmake --build build/release -j 4` (including all standalone
I/O benchmark targets)
- focused lifecycle/configuration/request tests: 104 test cases, 805
assertions
- focused cache tests: 48 test cases, 2,099,946 assertions

## Checklist

- [x] Read `CONTRIBUTING.md` and meet the PR reviewability requirements.
- [x] Cover cache, reactor, readahead, scan-manager, and parquet
behavior with tests.
- [x] Document backend, cache, and pinning configuration changes in code
and user documentation.
- [x] Update performance scripts, benchmarks, and agent documentation.

## Stack

This was layer 6 of the frugal native-I/O stack; the lower layers
(sirius-db#1773, sirius-db#1774, sirius-db#1776, and sirius-db#1772) are already merged. Supersedes sirius-db#1771:
`dev` was folded into `main` upstream, and this branch is now rebased
directly onto `main`'s tip as a normal, non-stacked PR.


----

GB300 benchmarks:

profile  lukewarm. hot           cold
---------------------------
q1           4.0062      2.7908
q2           1.0207      0.3670
q3           1.7495      1.1402
q4           1.9325      0.8274
q5           2.1787      1.1928
q6           0.9368      0.9093
q7           1.3715      1.4290
q8           2.4953      1.1644
q9           2.8493      3.0254
q10          2.5605      1.8375
q11          0.4259      0.3570
q12          1.2622      1.3134
q13          2.1456      0.9090
q14          1.1204      1.1168
q15          1.0552      0.9031
q16          0.4655      0.3556
q17          1.3752      1.2920
q18          2.6438      2.2206
q19          1.4832      1.4426
q20          1.2476      1.2262
q21          2.7803      2.7661
q22          0.3547      0.3392
-------------------------------
Total       37.4606     28.9252


cold 
Query         iter0
-------------------
q1           3.7907
q2           0.9496
q3           2.5600
q4           1.6931
q5           3.0860
q6           2.0780
q7           3.1990
q8           4.3953
q9           5.0900
q10          3.2508
q11          0.6380
q12          2.3308
q13          2.6227
q14          2.3591
q15          2.7282
q16          0.5409
q17          2.4587
q18          3.0535
q19          2.8955
q20          3.5231
q21          4.1316
q22          0.7383
-------------------
Total       58.2130

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Matthijs Brobbel <m1brobbel@gmail.com>
Bumps CUDA 13 env to latest 13.4 release.
…pare walk (sirius-db#1578)

# perf(scan): cache + fuse + parallelize the duckdb-native metadata
prepare walk

**Branch:** `pr/unpinned-metadata-walk-cache` → base: **`main`**.
Rebased onto current `main` (2026-09-29). The two new headers follow
sirius-db#1806's move to `src/op/scan/`, and the new sources are registered in
`src/op/CMakeLists.txt` and `cmake/sirius-test-sources.cmake` per sirius-db#1905.

## Background: what the "metadata prepare walk" is

When Sirius scans a table stored in DuckDB's native format, it first has
to plan the scan before touching any data. That planning step walks over
every **row group** (DuckDB stores a table as a sequence of row groups,
~122k rows each; a SF1000 `lineitem` has ~49k of them) and, for each
one, asks DuckDB for:

1. its geometry (where it starts, how many rows it holds),
2. the min/max statistics of every column that has a pushed-down filter,
so whole row groups that can't match can be skipped ("pruning"),
3. the max string length of every projected `VARCHAR` column, to make
sure no row group contains an oversized string that the GPU string
decoder can't handle (if one does, Sirius refuses the GPU path for that
scan).

We call this the **metadata prepare walk**. Today it is done from
scratch on every query, serially, on the query thread, while the
query-lifecycle slot is held. On a `lineitem`-sized table at SF1000 it
costs **57–59 ms warm and 1.33 s cold** per query.

Since sirius-db#1548 merged, tables whose data is already **pinned** in GPU
memory skip this walk entirely. This PR is for everything else:
**unpinned** tables, including the first query against a table before
its data is pinned. Those still have to do the walk, so this PR makes it
cheap.

## What changes

Three things, in order of how much of the work they remove:

### 1. Cache the per-table snapshot

Steps 1–3 above read data that only changes when the table's physical
layout changes (insert, delete, checkpoint). So we now capture it once
per table into an immutable **snapshot**: the row-group geometry plus a
copy of every row group's column statistics. Later queries against the
same table reuse the snapshot instead of asking DuckDB again, which
means no storage reads.

Getting this right under concurrent writers is where most of the care
went:

- The capture is taken in a single pass while holding DuckDB's
segment-tree lock, so it can't observe a half-committed table. If a
commit lands mid-capture anyway, the capture is detected as **torn**
(inconsistent) and retried rather than cached.
- On every reuse the snapshot is validated against the live table by
comparing row-group identities via `weak_ptr`. A `weak_ptr` can only
lock back to the exact object it was created from, so a row group that
was freed and a new one allocated at the same address is correctly seen
as different. (This is the "ABA-safe" property: it protects against the
"A was replaced by B, then something else that looks like A" bug.)
- If the current transaction has uncommitted appends to the table
(DuckDB's `LocalStorage`), the cache is bypassed for that query, because
the snapshot can't see those rows.

### 2. Cache the per-query-shape result on top of that

Steps 2 and 3 depend not just on the table but on **which columns are
projected and which filters are pushed down**. Two queries with the same
projected columns and the same filter predicates produce the same
pruning decisions and the same oversized-string verdict. So we
additionally cache the *outcome* of the walk, keyed by (projected column
set, pushed-down filter set).

- Filters are compared with `TableFilter::Equals`, restricted to filter
types whose comparison actually covers all their parameters, so we never
treat two different predicates as equal.
- Each cached result is tagged with the snapshot generation it was
computed from and is discarded the moment the snapshot rebuilds, so a
result can never be served against a table layout it wasn't computed
for.

A hit here reduces the whole prepare step to a single validity probe,
with no statistics checked at all.

### 3. Make the remaining walk faster when it does have to run

When neither cache hits (first query against a table, or a new
predicate), the walk still has to look at statistics. Previously that
was one full pass over all row groups **per filter column** plus one
**per varchar column**. Now it is a single fused pass over the row
groups that checks all columns at once, and that pass is parallelized
across row groups.

Parallelizing this is safe because `RowGroup::GetStatistics` takes its
own lock per row group and column and returns a self-contained copy, and
every `TableFilter::CheckStatistics` implementation is `const` and
read-only. Only `GetPartitionStats` has to stay serial (it touches
`LocalStorage` and the `ClientContext`). Workers write to disjoint
pre-indexed slots, and when a refusal is needed the result is reduced to
exactly the same (column, row group) the old serial passes would have
reported, so the observable behaviour is unchanged.

**Env knobs:** `SIRIUS_METADATA_WALK_THREADS` caps the worker count (`1`
= serial); `SIRIUS_DISABLE_NATIVE_METADATA_CACHE` restores the old
uncached walk.

## Key review point: why cache invalidation is probe-based rather than
counter-based

The tempting cheap approach is to key the cache on the table's last
commit id (`GetLastCommit`). That is **wrong**: `CHECKPOINT`
restructures row groups (compaction) without bumping the commit counter,
so a counter-keyed cache would serve stale row-group geometry after a
checkpoint. The identity probe described in section 1 is what actually
makes this safe. There is an explicit "rebuilds after checkpoint
compaction" test so that a future "optimization" back to counter keying
fails CI.

## Measured (SF1000 `lineitem`, 48,849 row groups, unpinned, warm;
`[walk_bench]` harness)

| query shape | before | cache disabled (fused + parallel only) | first
query (builds snapshot) | new predicate (snapshot hit) | repeat (result
hit) |
|---|---|---|---|---|---|
| q1-heavy (6 varchars) | 56.6 ms | 19.1 ms | 12.4 ms | 7.6 ms | **6.1
ms (−89%)** |
| q1 (2 varchars) | 30.0 ms | 14.8 ms | 8.5 ms | 6.9 ms | 5.5 ms |

Column meanings: "cache disabled" isolates section 3
(`SIRIUS_DISABLE_NATIVE_METADATA_CACHE=1`); "first query" is the first
walk after a table changes, which pays for capturing the snapshot; "new
predicate" is a query with an unseen filter against an unchanged table,
where section 3 runs over cached statistics with no storage reads;
"repeat" is an already-seen query shape served by section 2.

Cold walk: 1327 ms → 12.9 ms. Pin-served benchmark suite: neutral
(+0.6%, within noise). This is a latency win for unpinned tables and
first queries, not a suite mover.

## Tests

`test/cpp/scan/test_duckdb_native_metadata_cache.cpp`
(`[duckdb_native_metadata_cache]`, GPU-free), 16 cases:

- cached result equals the uncached walk; repeated walks hit; varied
filters are served from one snapshot;
- result cache hits on repeated query shapes, misses on new predicates,
and drops results when the snapshot rebuilds;
- invalidation on committed insert; stays valid across committed
deletes; rebuilds after checkpoint compaction; bypasses on
transaction-local appends; tables are keyed independently;
- varchar overflow refusal is preserved, and sub-limit varchars are
accepted;
- the fused pass reports the same refusal the old column-outer pass did;
the parallel walk is deterministic across worker counts;
- a torn snapshot is never served under concurrent commits; prepared
walks survive concurrent commits and settle exactly.

`test/cpp/scan/bench_metadata_walk.cpp` is the hidden `[walk_bench]`
harness used for the numbers above.

Run on the rebased branch (current `main` base):
`[duckdb_native_metadata_cache]` — 16/16 green (1392 assertions);
neighbor sweep `[scan],[filter]` — 395/395 green (462,504 assertions).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…nts (sirius-db#1847)

## Description

Compressed-materialization tests currently add six counters, six
recording methods, and a snapshot API to `SiriusContext`. This moves
those observations onto the query event publisher introduced in sirius-db#1776.
Scan, pinning, partition, and planner producers publish an activity and
count; only test subscribers aggregate counters. This follows up on the
[review discussion in
sirius-db#1765](sirius-db#1765 (comment))
and targets `dev` independently of that PR.

A subscriber `flush(timeout)` waits for callbacks for all deliveries
preceding its captured boundary. Pending event IDs are tracked
individually because producers can enqueue out of order. The wait
releases the publisher's routing lock and reports timeout, shutdown, or
delivery failure instead of allowing tests to assert against partial
observations. Existing before/after test helpers subscribe on their
first snapshot and flush before reading results, including assertions
that no activity occurred.

The change preserves each observation's existing units and producer
locations. It adds no configuration setting and changes no query
execution or partition-sizing decisions. Event tracking and counter
aggregation tests belong together here because asynchronous delivery
must be synchronized before replacing the existing synchronous counters.

## Checklist

- [x] Read CONTRIBUTING.md and keep this independent change scoped for
review.
- [x] Add event and synchronization tests; migrate existing
integration-test observation helpers.
- [x] Update compressed-materialization documentation.
- [x] No user-facing configuration changes.
- [x] Run the full Linux/CUDA build and affected integration tests.
…ltered decode, and key-type generalization (INT8..UINT64, DATE/TIMESTAMP, DECIMAL, STRING, nullable keys) (sirius-db#1557)

> **2026-09-29: rebased onto `main` at `b729ac0c` (the repo's default
branch was renamed from
> `dev`) and review round 2 addressed.** Head `8c131452`; previous tip
`b04aba21`. Two commits:
> `092db3d9` is the PR itself (rebased; rebase-only conflicts in
`CMakeLists.txt` -> main's
> per-directory source lists, the publisher's new publication-session
`try` block from sirius-db#1843, and
> the scan operator's `snapshot.entries()` loop), and `8c131452` holds
only the changes made in
> response to @joosthooz's seven threads, so the review delta is
isolated:
>
> - **Narrowed build keys are published uniformly** (thread 1, the
"silently dropping probes"
> one): the DATE-only cast and the `==` zone-map gate are replaced by
one domain rule
> (`membership_same_family`); zone-map bounds are always restored to the
recorded type, and
> membership filters are pushed only to bindings whose probe type the
domain can read
> (`membership_probe_compatible`, counted in
`bindings_skipped_incompatible_probe`). New tests:
> INT64 key from an INT16 build, DECIMAL64 key from a DECIMAL32 build,
compat gate.
> - **One source of truth for the carrier rule**:
`sirius::value_fits<To>(v)`
> (`helper/numeric_carrier_rule.hpp`, `__host__ __device__ constexpr`)
is used by the device
> `probe_key_convert` and the host `numeric_range_fits`;
`integer_storage_type` is the single
> timestamp->integer mapping (all units map; the narrowing helper
deliberately applies it to DATE
> only, documented); `column_values_fit` in `numeric_narrowing` replaces
the decimal-specific fit
> check; `narrow_domain_of(cudf::data_type)` is the one carrier-family
list.
> - **Device == host test**: every (rep, carrier) pair at every boundary
value vs `value_fits`.
> - **Dedup**: `compute_mask` prior-free forwarder defined once in
`sirius_mask_applicable`;
> constructor prologue shared as `prepare_membership_build`; probe
kernel prologue/launch shared
> as `membership_probe_functor` / `run_membership_probe`, each filter
contributes only its lookup.
>
> Kernel count unchanged (19 per filter kind). Verified at `8c131452`:
release + clang-debug
> builds clean; `[dynamic_filter]` 306/306, `[fused_scan_filter]` 27/27,
`[numeric_narrowing]`
> 14/14; full `make test` 3677/3680 with the 3 failures being a
since-fixed assertion in the new
> narrowing test (tag re-run green) and two environment/flaky cases
unrelated to this PR
> (`sirius_config forces the sirius backend for multi-GPU` on a 1-GPU
host; `test_explicit_eviction`
> from sirius-db#1908 passes in isolation). `pre-commit run -a` clean.
>
> **2026-09-21: rebased onto `dev` at `b1909fcc` and squashed to one
commit** (new head
> `b04aba21`; previous tip `743eed39`, incl. the review-cleanup commit,
is preserved on the
> fork as `pr/carrier-native-df-probes-rebased-dev`'s ancestor and in
the reflog). Squashed
> because sirius-db#1817 (`rmm::cuda_stream_view` -> `cuda::stream_ref`) rewrote
the signatures every one
> of the 8 commits touched, so a per-commit rebase conflicted 8 times
over. Adaptations to dev:
>
> - **sirius-db#1806 header split**: `src/include/` is gone; the new files now
live beside their sources
> (`src/cuda/dynamic_filter_probe.cuh`,
`src/op/dynamic_filter/dynamic_filter_key_domain.hpp`)
>   and the `selection.hpp` edit followed that header to
> `src/compression/simpatico_codegen/src/codegen/selection/`. Include
paths are unchanged.
> - **sirius-db#1817 stream migration**: all PR code and tests use
`::cuda::stream_ref` (`.get()`,
>   `.sync()`); no `cuda_stream_view` remains in PR-touched files.
> - `restore_probe_to` call sites that dev still had are gone with it;
no references remain.
>
> Verified at `b04aba21`: `pixi run make` and `pixi run make
clang-debug` both clean,
> `[dynamic_filter]` 277/277, `[fused_scan_filter]` 27/27, `make test`
3367/3367 on an idle GPU,
> `pre-commit run -a` clean.
>
> **2026-09-17: scope extended.** Six commits (`d219a326`..`b352168c`)
generalize the
> membership probes to every fixed-width integer, DATE/TIMESTAMP,
DECIMAL, STRING and nullable
> keys via a key-domain dispatch layer — see **section 3** below. Head
moved `80801eb1` ->
> `b352168c`; full `make test` 3293/3293.
>
> **2026-09-16: rebased onto `dev` at `f54e6a10`** (new head `80801eb1`;
the previous tip
> `30fe3dd9` is preserved on the fork's reflog and referenced here).
Still a single commit.
> Three files conflicted:
>
> - `CMakeLists.txt` — test-source list; kept both dev's new
iceberg/puffin tests and
>   `test_fused_membership_mask.cpp`.
> - `src/cuda/sirius_dynamic_bloom_filter.cu` — include block only. Dev
(sirius-db#1809) dropped the
> custom `sirius_bloom_policy` in favour of
`cuco::default_filter_policy`, so the
> `<cuco/hash_functions.cuh>` include this PR carried is dead and was
dropped too; the
>   `bloom_contains` Bulk kernel is unchanged.
> - `src/compression/simpatico_codegen/src/simpatico_codegen.cpp` — dev
landed the T9
> positional keep mask (sirius-db#1556) as an extra AND source appended *after*
wave 1, while this
> PR moves the combine *before* the sequential membership cascade.
Resolved by making the
> keep mask a cascade prior: it is uploaded on `s0` before the combine
and counts as
> "another source", so a membership-only chunk *with* a keep mask now
takes the sequential
> pruned path. With no static range/bool8 strip nothing else writes
`combined`, so the keep
> mask lands there directly instead of a separate strip. The "every
probe declined" fallback
> is unchanged and still correct on that path (`keep_mask_applied` stays
false, the caller
> applies the mask itself). A new case in
`test_fused_membership_mask.cpp` covers it.
>
> The "Deliberately not carried over" paragraph below is therefore
partly stale: the T9
> keep-mask **is** now a prior source; pair (col-vs-col) filters remain
out.
>
> Verified on this rebase: `[dynamic_filter]` 241/241,
`[fused_scan_filter]` 22/22 (5 fused
> membership cases), `make test` 3257/3257 — 36 `gpu_execution iceberg`
cases failed on the
> first pass only because the `iceberg` DuckDB extension was not
provisioned on the box;
> all 36 pass after `INSTALL avro; INSTALL iceberg`. `pre-commit run -a`
clean.

> **Rebased onto `dev`.** This PR previously sat on the closed
fused-scan-filter
> (sirius-db#1391) + late-mat (sirius-db#1409) stack — the "wave-2 lineage" the old
description
> referred to. Both were reimplemented and merged as sirius-db#1474 and sirius-db#1524, so
the
> branch is now a **single commit off `dev`** with no base-branch diff
> pollution. The old wave-2 tip is preserved at `a9ac3d88`.

## 1. Carrier-native probes

`compute_mask` on all three membership filters required the probe
column's type
to equal the filter set's type. Compressed materialization stores a
bounded
`BIGINT`/`INTEGER` key in the narrowest fitting signed carrier, so a
decode-time
probe sees `INT32` against an `INT64`-built set. The probes now accept
any
signed-integer carrier (`INT8`..`INT64`) and convert per element
in-kernel. A
value the key domain cannot represent is a definite non-member, so the
no-false-negative contract holds.

### This replaces `restore_probe_to`, and the two were measured against
each other

sirius-db#1555 fixed the same defect by materializing a widened copy of the probe
column
before the lookup. The two are **independent fixes for the same problem,
derived
two weeks apart** — neither the wave-2 base nor T2's own branch had
`restore_probe_to`; it appeared during sirius-db#1555's own rebase. Merge order
rather
than merit picked the winner, so I A/B'd them.

Pooled allocator (production hands these calls a pooled `mr`; an
unpooled one
charges the restore a driver allocation the real path never pays), both
arms
interleaved in one process, median of 15 warmed reps:

| filter | rows | filters/col | A: restore | B: in-kernel | B/A |
|---|---|---|---|---|---|
| hash IN-list | 1M | 1 | 23.6 µs | 19.0 µs | **0.81** |
| hash IN-list | 8M | 1 | 92.1 | 67.7 | **0.74** |
| hash IN-list | 8M | 3 | 256.4 | 200.5 | **0.78** |
| Bloom | 1M | 1 | 19.6 | 16.2 | **0.82** |
| Bloom | 8M | 1 | 67.0 | 46.1 | **0.69** |
| Bloom | 8M | 3 | 178.7 | 117.9 | **0.66** |

In-kernel is faster at **every point measured**, across two independent
runs,
and both arms produce identical masks (asserted in the harness). The
margin
widens with filters per column because `restore_probe_to` sits *inside*
`compute_mask` and so re-materializes the same column once per filter —
which
the post-decode cascade in `dynamic_filter_merge` does routinely, on the
default
path, not just under the experimental gate.

Read the pooled numbers as a **lower bound**: the benchmark pool is
fresh,
dedicated and uncontended, so production's allocation costs at least
this much.

Scale: ~3 µs per million rows, so ~18 ms for a full lineitem probe pass
at
SF1000. Real and consistent, but modest in absolute terms — this is not
a
headline win, and nothing here recovers the number the old description
claimed.

`restore_probe_to` had no other callers and is removed with its helpers.

## 2. Mask-aware probing

`sirius_mask_applicable` gains a prior-keep-mask `compute_mask` overload
(packed
1 bit/row); dead rows skip the set/Bloom lookup. The wave-1 orchestrator
runs
membership probes sequentially on `s0` after combining the other mask
sources,
handing the combined words to each probe as its prior and folding each
result
back in, so later probes see earlier survivors. Membership-only chunks
keep the
concurrent, prior-free path. The prior is a hint only: ignoring it is
sound
because the caller ANDs with that same mask.

**This half is unmeasured and is the open question on the PR.** It
trades
concurrency for pruning: a probe that previously ran round-robin on its
own pool
stream now serializes on `s0` behind the combine, and only the lookup is
pruned
— the full-width key decode is still paid. At high static keep-rate that
may not
pay for the lost overlap. **Reasonable to split out** if you'd rather
land item 1
on its own measurement and take item 2 separately.

## 3. Key-type generalization: membership probes beyond `INT32`/`INT64`
(added 2026-09-17, commits `d219a326`..`b352168c`)

Item 1 fixed the *carrier* mismatch for integer keys. This section
removes the *type* gate
entirely. All three membership filters (hash IN-list, small IN-list,
Bloom) previously accepted
only `INT32`/`INT64` keys with `null_count()==0`; any other equality key
reached the publisher and
was silently assigned no membership filter (`choose_membership_filter`
-> `none`). Type gating
lived in three places, none of them the planner: the filters'
`supports()`, the publisher's
build-column type-equality check, and the join-edge route gate in the
publish plan.

### Dispatch design

Two orthogonal, explicitly enumerated axes instead of a
`cudf::type_dispatcher` over the
(key, probe) type pair (which would instantiate hundreds of kernels):

- **Key rep** — the device element type the set / needles / Bloom are
built over:
  `int32_t`, `int64_t`, `uint32_t`, `uint64_t`. New host-only header
`src/include/op/dynamic_filter/dynamic_filter_key_domain.hpp`:
`membership_key_family`
`{signed_int, unsigned_int, date_days, timestamp, decimal,
string_hash}`,
`classify_membership_key`, `membership_key_supported`,
`membership_probe_compatible`,
`membership_build_fits_rep`. These replace every `INT32|INT64` literal
in the three
  `supports()`, the publisher, `dynamic_filter_publish_plan.cpp`, and
  `direct_route_admissible`.
- **Probe adapter** — a device functor chosen on the host from `(domain,
probe.type())`:
`integral_probe_adapter<ProbeRep, KeyRep>` (generalizes item 1's
`probe_key_convert`;
signedness-aware, range-checked, sentinel-aware), and
`string_hash_adapter`. Kernels
(`set_contains`, `small_in_list_scan`, `bloom_contains`) take the
adapter, a `probe_validity`
bitmask reader and the optional prior keep-mask, and emit a
**non-nullable** `BOOL8` mask.

Instantiation budget: **19 probe kernels per filter kind** (was 8 on
dev, 16 for the integer
scaffold, +2 for `__int128` decimal carriers, +1 string). Compile time
per `.cu` unchanged
within noise (~30 s each, measured cache-busted).

### What each family does

| Family | Rep | Build side | Probe side | Notes |
|---|---|---|---|---|
| `INT8`/`INT16`/`INT32`/`INT64`, `UINT8`..`UINT64` | i32/i64/u32/u64 |
any signed/unsigned integer column; the publisher now **accepts a build
column arriving at a narrower carrier** when `can_restore_to` holds
(previously the key was skipped, `keys_skipped_type_mismatch`) | any
carrier of the same signedness family, converted in-kernel;
out-of-domain value = definite non-member | sentinel is `min()` for
signed, `max()` for unsigned reps |
| `TIMESTAMP_DAYS` (DATE), `TIMESTAMP_{S,MS,US,NS}` | i32 / i64 | reads
the underlying rep; a DATE build column arriving as `INT8`/`INT16` is
**restored to `TIMESTAMP_DAYS`** in the publisher (a carrier-typed set
would decline the native probe the post-decode cascade emits) | same
unit, or `INT8`/`INT16`/`INT32` carriers for DATE; mixed units decline
(unreachable anyway: DuckDB inserts a cast and a cast blocks the scan
route) | zone maps already supported these and serve as the test oracle
|
| `DECIMAL32`/`DECIMAL64`/`DECIMAL128` | i32 / i64 / i64-if-fits |
`DECIMAL128` builds are range-probed (`compute_exact_numeric_range`); if
the unscaled range fits `int64` the set is `int64`, else **all three
membership filters decline** (zone map still publishes) | same scale
required; `DECIMAL32`/`64`/`128` carriers all accepted, `__int128`
probes range-checked into `int64` | no `__int128` Bloom (deliberately) |
| `STRING` | u64 | `cudf::hashing::xxhash_64` fingerprints materialized
once per build | `XXHash_64<string_view>` computed **in-kernel** over a
`column_device_view`, no materialized copy; `DICTIONARY32` declines (the
fused decode reconstructs dictionary/str_split carriers to `STRING`
first) | both IN-lists become **no-false-negatives** for strings
(fingerprint collisions keep rows); the sentinel `UINT64_MAX` needs no
remap — cuco's insert of its empty key is a no-op and the probe keeps
rows whose fingerprint equals it |
| **nullable keys** (any family) | — | `supports()` no longer requires
`null_count()==0`; null build keys are compacted out (`drop_nulls`, as
Bloom already did); the small IN-list tier gates on *valid* rows; the
publisher sizes/tiers on `valid_rows` | null probe rows emit `false` via
the `probe_validity` prologue; `copy_bitmask` tails removed | exact, not
approximate: null-safe keys are never admitted and the join runs
`null_equality::UNEQUAL`, so a NULL never matches on either side |

Not done, on purpose: `FLOAT32/64` (no benchmark uses float keys; would
need -0.0/NaN
canonicalization), `BOOL8`, `LIST`/`STRUCT`, `HUGEINT` (its lossy
`INT64` mapping is a separate
bug).

### Where the expected TPC-H wins turned out not to exist

TPC-H has **only `INTEGER`/`BIGINT` join keys**, so this section is
inert on TPC-H beyond the
narrow-build-carrier acceptance and nullable-key handling. In
particular:

- q2's `ps_supplycost = min(...)` is planned as a `FILTER` above a
`DELIM_JOIN`, not as a join
  condition — there is no DECIMAL join key to admit.
- q15's `total_revenue = (subquery)` *is* a DECIMAL128 join condition,
but its build side is
`AGGREGATE max(...)` over a `CTE_SCAN`, which `build_relation_is_opaque`
/
`build_subtree_is_filtering` do not recognize as evidence. That is the
evidence policy,
  upstream of admission, and is untouched here.

The beneficiaries are TPC-DS. A survey at SF10 on the integer-only
scaffold found that of the
21 queries with STRING/DATE/DECIMAL/nullable join keys, only ~7
currently run on the GPU at
all (the rest fall back on UNION / WINDOW / DISTINCT / mixed-inequality
joins, or crash: q10
aborts on a MARK-join mode assertion). Those seven *do* show the key
being declined today:
q83 (`d_date` -> `none`), q59 (`s_store_id` -> `none`), q30/q81/q64/q69
(nullable INT keys
downgraded IN-list -> Bloom). No TPC-DS timing is claimed in this PR.

**SF1000 TPC-H QphH, same recipe as August, no patched libcudf:** dev
`f54e6a10` 7.79M
(Power 10.15M / Tput 5.98M); this branch at `80801eb1` (items 1+2 only)
median of 3 clean runs
7.95M (Power 10.45M / Tput 6.04M) — Power +3.3%, inside the ±4.5%
run-to-run band. Firmer than
the score: dev hit GPU-OOM retry-cap CPU fallbacks in 3 of 4 throughput
phases, this branch in
0 of 3. Per-query the membership-filter consumers moved most (q17 −11%,
q8 −9%, q7/q20 −7%,
q19 −6%). The type-generalization commits were not scored separately:
they cannot change any
TPC-H plan.

**Re-measured 2026-09-21 at the rebased head `b04aba21` vs dev
`b1909fcc`** (4 clean interleaved
runs each, medians): Power 10.49M vs 10.19M (+2.9%), Throughput 6.02M vs
5.99M (+0.5%), QphH
**7.94M vs 7.83M (+1.5%)**, stream-0 7.85 s vs 8.07 s. All four PR Power
values sit above all four
dev values; q19 −17%, q17 −11%, q8 −7%, q20 −6%, q7 −5%; nothing slower
than dev by >0.3%.
Side finding, not caused by this PR: on this dev base the 7-stream
throughput phase hits
GPU-OOM retry-cap CPU fallbacks in 6 of 10 dev attempts (2 of 6 with the
PR; 0 of 1 on the
f54e6a1-based control), which then either kills the streams on DuckDB's
32 GB limit or spills
hundreds of GB to disk. Power phases are unaffected.

### Tests for this section

`[dynamic_filter]` 277 cases / `[fused_scan_filter]` 27 cases; full
`make test` 3293/3293 on an
idle GPU at `b352168c`. New coverage, per family and for all three
filters, with and without a
prior mask, against a host oracle and (for DATE/DECIMAL) the lowered
zone-map AST as an
independent oracle: cross-carrier identity `INT8..INT64` x
`{INT32,INT64}` keys vs. the old
path; `INT8`/`INT16`/`UINT*` keys and unsigned sentinel semantics;
`TIMESTAMP_DAYS` x
`{DAYS, INT8, INT16, INT32}` probes and unit-mismatch declines;
`DECIMAL{32,64,128}` x
`DECIMAL{32,64,128}` probes, scale mismatch declines, DECIMAL128 fit ->
int64 set vs. overflow
-> decline; STRING host XXH64 reference pinned against
`cudf::hashing::xxhash_64` over a corpus
covering every XXH64 tail path (empty, 1 KiB, UTF-8, embedded NUL);
nullable builds x nullable
probes x 4 carriers, sliced-probe offsets, all-null build/probe;
publisher accepting
`INT8`/`INT16` build carriers, `DECIMAL32`-carrier for a `DECIMAL64`
key, `INT16`-carrier DATE;
fused decode of `INT8`-carrier, `INT16`-carrier-DATE,
`DECIMAL32`-carrier and dictionary-`STRING`
chunks against the widened sets. `docs/super-sirius/dynamic-filters.md`
gains a "Key types"
section with this table.

## The original's numbers do not transfer

The old description claimed SF1000 Power **+7.2%**, q19 **−52%**. Those
were
measured against a stack where a declining probe *threw* and failed the
entire
filtered decode. sirius-db#1524 changed that to a stand-down to the AND identity,
and
sirius-db#1555 removed the decline for narrowed carriers altogether. **Treat the
old
numbers as void.**

## Notes

- **The multi-probe cascade is unreachable in a default build.**
`SIRIUS_EXP_FUSED_SCAN_MAX_MEMBER` defaults to `1`, so the fold-back
never
executes; the new test arms it to 2. Worth deciding whether that default
moves.
- **Small-in-list was not benchmarked** — the A/B covered hash IN-list
and Bloom.
- Item 2's decode-time path sits behind `SIRIUS_EXP_FUSED_SCAN_FILTER`
(off
unless set). Item 1 is on the default path via the post-decode cascade.

## Deliberately not carried over

`dev`'s reimplementation dropped pair (col-vs-col) filters and the T9
visibility
keep-mask (still open as sirius-db#1556), both of which the original also primed
priors
from. Priors here are ranges and dictionary-answered equalities only. A
declining probe in the sequential cascade leaves the mask untouched,
matching
`dev`'s all-ones stand-down rather than the original's throw.

## Related, not fixed here

The build side has the mirror defect with a worse failure mode:
`dynamic_filter_publisher.cpp` compares the runtime build column against
the
plan's recorded logical type and **skips the key entirely** on a
mismatch
(`keys_skipped_type_mismatch`), so a narrowed build key publishes no
filter at
all — probe-side tolerance cannot help there. Already instrumented via
`dynamic_filter_stats`; worth checking whether that counter fires.

## Tests

- `test/cpp/operator/test_dynamic_filter_probe.cpp`
(`[dynamic_filter][probe]`,
  8 cases): heterogeneous carriers, decimal/date/float refusal, sentinel
conservation, prior-mask all-dead/all-live/patterned, null propagation.
- `test/cpp/scan/test_fused_membership_mask.cpp` (`[fused_scan_filter]`,
4 cases): prior handed alongside a static range, prior-free
membership-only,
a two-probe cascade proving the fold-back, all-dead prior, INT32 carrier
vs
  INT64 set.

Against current `dev`: `[probe]` 8/8, `[fused_scan_filter]` 4/4,
`[dynamic_filter]` 240/240, **full suite 3035/3035** (32,788,590
assertions).
`pre-commit` clean. The A/B harness was scaffolding and is not in the
diff.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…arrocks (sirius-db#1911)

Bumps [thiserror](https://github.com/dtolnay/thiserror) from 2.0.20 to
2.0.21.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/dtolnay/thiserror/releases">thiserror's
releases</a>.</em></p>
<blockquote>
<h2>2.0.21</h2>
<ul>
<li>Fix parsing of generic unit variants in display expressions (<a
href="https://redirect.github.com/dtolnay/thiserror/issues/459">#459</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/dtolnay/thiserror/commit/b1827ee06f81a7d7f954e676771e19a97f9f8e3b"><code>b1827ee</code></a>
Release 2.0.21</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/58037b575a2a5543d0d54a49360554325b49f55a"><code>58037b5</code></a>
Merge pull request <a
href="https://redirect.github.com/dtolnay/thiserror/issues/459">#459</a>
from dtolnay/turbofish</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/f82a0cf8f06c71263c2caa26e759f337d2778153"><code>f82a0cf</code></a>
Keep track of nested turbofish depth</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/72ea49262d6412ccbc08f1696dda8626409e657d"><code>72ea492</code></a>
Raise required compiler to Rust 1.77</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/72eea0d4ddb17ff9fa873aa74469de850021b22e"><code>72eea0d</code></a>
Resolve io_other_error clippy lint in tests</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/07f09a2ec934df58508ce4fa5e7fbf27cca552c2"><code>07f09a2</code></a>
Raise required compiler to Rust 1.74</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/2715388e8cc5787b3643251f7a7a637e89376688"><code>2715388</code></a>
Update ui test suite to nightly-2026-09-22</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/5a306c7d0a8588caaaaa6a7567aeb25c1c10719b"><code>5a306c7</code></a>
Update ui test suite to nightly-2026-09-05</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/ef9383b37d7c96b01e10df35d3e9e313c7b98b5a"><code>ef9383b</code></a>
Update ui test suite to nightly-2026-08-22</li>
<li><a
href="https://github.com/dtolnay/thiserror/commit/8336b8407fbfe69177c504cdbc379db20cf6f131"><code>8336b84</code></a>
Update ui tests for version 2.0.20</li>
<li>See full diff in <a
href="https://github.com/dtolnay/thiserror/compare/2.0.20...2.0.21">compare
view</a></li>
</ul>
</details>
<br />


[![Dependabot compatibility
score](https://dependabot-badges.githubapp.com/badges/compatibility_score?dependency-name=thiserror&package-manager=cargo&previous-version=2.0.20&new-version=2.0.21)](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores)

Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore this major version` will close this PR and stop
Dependabot creating any more for this major version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this minor version` will close this PR and stop
Dependabot creating any more for this minor version (unless you reopen
the PR or upgrade to it yourself)
- `@dependabot ignore this dependency` will close this PR and stop
Dependabot creating any more for this dependency (unless you reopen the
PR or upgrade to it yourself)


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…an_manager_state) (sirius-db#1366)

This is PR 3 from a stack of PRs (it goes on top of
sirius-db#1364 and
sirius-db#1365) which together will make
Sirius internals be able to handle Concurrent query execution. There is
more to come.

## Summary 
This PR addresses the scan_manager. It primarily takes all its internal
structures that are query specific and puts them into a
query_scan_manager_state. Then the scan_manager can hold a map of these
states.

## Background

The scan manager used to hold one query's worth of state in bare members
(_providers_by_op, _scan_op_order, _metadata_processor, _dispatcher,
_pending_mvcc_mask_jobs, _pending_insert_delta_jobs). reset() wiped all
of
them. This was fine for a single serial query but breaks under
concurrent
queries

## What changed

### 1. query_scan_manager_state

All per-query state is now grouped in a new inner struct,
query_scan_manager_state, held as a shared_ptr in a map keyed by
query_id:

std::map<query_id_t, std::shared_ptr<query_scan_manager_state>>
_query_states;
    mutable std::mutex _query_states_mutex;

State is published to the map AFTER prepare_for_query fully builds it,
so
a failed prepare never leaves a half-built entry visible. The shared_ptr
is
resolved under the lock and used outside it, so an erase racing a reader
cannot pull the state out from under the reader.

### 2. reset(query_id) / reset_all()

reset() now takes a query_id and only tears down that query's
dispatcher,
coalescer, and providers. Other queries' state is not touched.

reset_all() is the teardown-only path used by stop() and the
failed-query
backstop.

sirius_context::run_mandatory_cleanup now calls reset(query_id) instead
of
reset(). sirius_context::drop_query_runtime_state_best_effort (the error
backstop) now also calls reset(query_id): prepare_for_query no longer
performs a global reset.

### 3. Shared-ownership pin table

_pinned_entries changed from:
    std::unordered_map<std::string, pinned_entry>
to:
    std::unordered_map<std::string, std::shared_ptr<pinned_entry>>

A query that matched a pinned entry now co-owns it for its whole
duration
through the shared_ptr the cached_databatch_provider holds. A concurrent
UNPIN on another connection only drops the map slot; the data stays
alive
until the last serving provider releases it.

find_pinned_entry_for_duckdb_table changed from returning a raw
pinned_entry const* to a shared_ptr<const pinned_entry> (owning), so
plan-time MVCC guards that read the entry hold it alive across the
check.

make_provider_for_pinned_entry signature changed from
pinned_entry const& to std::shared_ptr<pinned_entry const> for the same
reason.

All pin-table mutations (insert, attach_mvcc_metadata, remove,
visit_pinned_entries) are now guarded by _pinned_entries_mutex.

try_match_cached_entry snapshots the pin table under
_pinned_entries_mutex
and matches outside the lock: the match path (can_serve_with_columns,
validate, zone-map plan building, MVCC branch) is too slow to hold the
lock across.

### 4. Thread pool sizing

The pool is now sized num_threads + k_max_concurrent_queries (was
num_threads + 1). Each query's coalescer parks exactly one sequencer
task
in queue.wait_dequeue, only unblocked by that query's own split_provider
tasks running on the same pool. With Q concurrent queries, Q threads are
parked in sequencers at all times, so the old +1 sizing would deadlock
at
2 concurrent queries. k_max_concurrent_queries is left at 1 (matching
the
old behavior for single-query runs) with a TODO to promote it to a real
config option.

### 5. Known gap documented

The prefetch cache epoch is a global counter
(prefetching_cache::_ticker),
so a second query starting bumps the epoch and demotes the first query's
prefetched-but-unconsumed chunks to tier-0 in the eviction order.
Performance only (not correctness: a live reader holds pin > 0, blocking
mark_evicting). The fix belongs in prefetching_cache, not here. A
comment
now calls this out clearly.

## Changed interfaces

  sirius_scan_manager::reset()     -> reset(query_id) + reset_all()
  sirius_scan_manager::find_pinned_entry_for_duckdb_table()
                                   raw pointer -> shared_ptr (owning)
  make_provider_for_pinned_entry() pinned_entry const& -> shared_ptr
  SiriusContext::run_mandatory_cleanup
scan_manager_->reset() -> reset(query_id)
drop_query_runtime_state_best_effort: scan_manager_->reset(query_id)
added

## New test: test_scan_manager_query_state.cpp

  - two queries register independently and reset drops only one
      Pins that tearing A's dispatcher down does not kill B's sequencer,
      and B's split connector still yields splits after A resets.
  - concurrent queries do not collide on operator id
Both queries get operator id 0. Asserts each operator drains splits
      for its own file only.
  - reset is a no-op for unknown and already-reset queries
The failure backstop calls reset(query_id) unconditionally on every
      code path; an unknown id must not throw.
  - reset_all drops every query and stop() still returns
      A sequencer parked on a dequeue with nothing to feed it would hang
      stop() indefinitely.
  - a query with no GPU scan operators registers nothing
The matching reset is a harmless no-op rather than touching the map.

## Files changed

  src/include/scan_manager/sirius_scan_manager.hpp   main API changes
  src/scan_manager/sirius_scan_manager.cpp           implementation
src/sirius_context.cpp per-query reset calls
  src/planner/sirius_plan_get.cpp                    owning shared_ptr
  test/cpp/scan_manager/test_scan_manager_query_state.cpp  (new)
test/cpp/scan_manager/test_cached_serving_hardening.cpp shared_ptr entry
  test/cpp/scan_manager/test_s3_routing_cutover.cpp  distinct query ids
  CMakeLists.txt
# Conflicts:
#	CMakeLists.txt
#	pixi.lock
#	src/creator/task_creator.cpp
#	src/creator/task_creator.hpp
#	src/cuda/tae/tae_decode_kernels.hpp
#	src/cudf/cudf_compat.hpp
#	src/data/host_tae_representation.hpp
#	src/data/host_tae_representation_converters.hpp
#	src/data/sirius_converter_registry.hpp
#	src/embedding/buffer_budget.hpp
#	src/embedding/control.hpp
#	src/embedding/execution_interrupted.hpp
#	src/embedding/execution_stats.hpp
#	src/embedding/input.hpp
#	src/embedding/native_gpu.hpp
#	src/embedding/plan_bindings.hpp
#	src/embedding/result.hpp
#	src/embedding/result_codec.hpp
#	src/embedding/result_gpu.hpp
#	src/embedding/source_wakeup.hpp
#	src/embedding/tae_demand.hpp
#	src/embedding/tae_read.hpp
#	src/embedding/terminal_admission.hpp
#	src/execution/sirius_execution_evidence.hpp
#	src/include/op/scan/gpu_ingestible_types.hpp
#	src/op/scan/tae_gpu_ingestible.hpp
#	src/op/scan/tae_scan_plan.hpp
#	src/pipeline/gpu_pipeline_executor.cpp
#	src/pipeline/gpu_pipeline_task.cpp
#	src/pipeline/gpu_stream_quiescence_error.hpp
#	src/pipeline/sirius_pipeline.cpp
#	src/planner/sirius_physical_plan_generator.cpp
#	src/sirius_c.h
#	src/sirius_ffi.cpp
#	src/tae/tae_format.hpp
…-dev-merge

# Conflicts:
#	pixi.lock
#	src/scan_manager/sirius_scan_manager.cpp
#	src/scan_manager/sirius_scan_manager.hpp
@aunjgr
aunjgr marked this pull request as ready for review September 29, 2026 10:53
@aunjgr
aunjgr merged commit e2e2f08 into matrixorigin:upstream-dev-merge Sep 29, 2026
14 of 21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.