Skip to content

Add optimizers for RPQ-matrix algortihm - #23

Merged
suvorovrain merged 26 commits into
mainfrom
rain/metadata
Oct 8, 2026
Merged

suvorovrain merged 26 commits into
mainfrom
rain/metadata

Conversation

@suvorovrain

@suvorovrain suvorovrain commented Jul 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR introduces cost-based optimization for matrix-based regular path queries (RPQs), supported by reusable graph metadata, additional GraphBLAS operations, and configurable benchmarking.

The optimizer uses egg to explore equivalent query plans and selects a plan using the chosen cost model. Optimization is performed during query preparation; the selected plan is subsequently executed on the full graph by the existing LAGraph-based evaluator.

Cost-based query optimization

The RPQMatrix implementation is reorganized into separate modules for query translation, rewrite rules, optimization, cost estimation, statistics, sampling, and execution.

The optimizer supports the following CLI strategies:

  • none — Executes the translated expression without e-graph optimization.
  • join — Uses a join-inspired estimate of multiplication work and scalar matrix statistics.
  • metaac — Uses density-based probabilistic estimates of intermediate result sizes.
  • hybrid — Combines join-inspired multiplication costs with MetaAC cardinality estimates.
  • mnc — Uses row and column count vectors to account for nonuniform matrix structure.
  • pang-hybrid — Combines scalar estimates with a Pang-inspired model of Kleene closure growth.
  • sampling — Estimates intermediate cardinalities by executing operations on sampled submatrices.

The rewrite rules cover reassociation of concatenation, distribution and factoring of alternatives, and replacement of expressions containing Kleene closure with specialized left- and right-closure operators. The rule set is initialized once and reused.

Cost models estimate both operation work and intermediate matrix characteristics. These estimates guide plan selection; they are not predictions of execution time in milliseconds.

MNC

The MNC model maintains row and column element counts, supplemented by singleton-related structural statistics. It uses these vectors to estimate multiplication and union cardinalities and propagates estimated statistics to parent expressions.

The implementation adapts MNC to Boolean RPQ relations rather than reproducing the original method in full. Intermediate counters may contain fractional estimates, and closure operators use fallback heuristics rather than a dedicated MNC closure model.

Pang-hybrid

The Pang-inspired model introduces an explicit estimate of closure growth instead of relying solely on a constant closure penalty. It simulates the growth of successive relation powers, estimates overlap with the accumulated result, and accounts for the work of closure steps.

The model uses scalar statistics rather than explicit vertex-domain intersections or Monte Carlo graph traversals. Its iteration limits apply only to cost estimation and do not restrict the paths considered during actual query execution.

Sampling

The sampling model selects a shared vertex subset for the labels used by a query and extracts the corresponding induced-subgraph matrices. Named query endpoints are explicitly included in the sample.

Boolean multiplication, union, and closure are performed on these smaller matrices during cost evaluation. The observed cardinalities and nonempty row and column counts are extrapolated to the full graph and combined with the analytical cost model.

Sampling is configurable through:

  • RPQ_SAMPLE_PERCENT: sampled vertex percentage, defaulting to 1%.
  • RPQ_SAMPLE_SEED: reproducible random seed.
  • RPQ_SAMPLE_MAX_STAR_ITERS: iteration limit for sampled closure evaluation.

This is an approximate estimation technique: paths through excluded vertices are absent from the sample. Analytical estimates are retained when a sampled result is unavailable, an incomplete sample produces no matches, or closure evaluation has not converged. Sampling affects plan selection only; final answers are computed on the full graph.

Graph metadata and storage

The in-memory graph representation is extended with per-label matrix metadata:

  • Matrix dimension.
  • Number of nonzero elements.
  • Number of nonempty rows and columns.
  • Optional row and column count vectors.
  • Optional extended statistics required by MNC.

The MatrixMarket loader computes and stores these statistics for reuse during optimization. Statistics collection is configurable through MatrixStatsMode; the CLI requests extended statistics for MNC without imposing that additional work on other strategies.

The optimizer can reuse precomputed statistics or construct them on demand when the graph backend does not provide them. Cached label statistics are associated with their underlying graph matrices to avoid reusing statistics from a different graph.

The MatrixMarket loader also maintains CSR and CSC matrix representations. Queries with a variable subject and a fixed object request CSC storage; other endpoint combinations request CSR. The graph interface provides a storage-aware lookup with a fallback for backends that do not maintain both orientations.

Fixed endpoints are represented by diagonal selector matrices and remain part of the expression processed by the optimizer.

Native GraphBLAS/LAGraph integration

The LAGraph dependency and Rust bindings are updated to support:

  • CSR/CSC orientation selection and orientation-aware matrix duplication.
  • Extraction of sampled submatrices.
  • Boolean operations and closure evaluation on sampled matrices.
  • Row and column count-vector construction.
  • Extended MNC statistics and count-vector estimation operations.

The native build script now tracks relevant LAGraph source, header, and CMake files so changes to the bundled dependency trigger a rebuild.

Benchmarking

This PR adds a fixed-run benchmark mode alongside Criterion-based benchmarking.

The fixed-run mode supports configurable warm-up and measured run counts and records two timing scopes separately:

  • Total: query preparation and execution.
  • FFI-only: execution of a prepared query, excluding preparation from the timed interval.

These scopes are measured in separate loops, not as components of the same execution. FFI-only timing includes the execution wrapper and result handling performed by execute, rather than isolating a single native call.

Aggregate statistics are written to the main result file, while individual measured durations are saved in a companion .runs.json file. Benchmark metadata records the selected optimizer and measurement configuration. Checkpoint validation also includes the optimizer and benchmark settings to reject incompatible resume configurations.

@suvorovrain
suvorovrain merged commit bb13874 into main Oct 8, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant