Repository navigation
Add optimizers for RPQ-matrix algortihm - #23
Merged
Merged
Conversation
suvorovrain
force-pushed
the
rain/metadata
branch
from
July 18, 2026 12:31
6969060 to
9782423
Compare
suvorovrain
force-pushed
the
rain/metadata
branch
from
September 26, 2026 13:22
e9e1092 to
3e3eb3c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR introduces cost-based optimization for matrix-based regular path queries (RPQs), supported by reusable graph metadata, additional GraphBLAS operations, and configurable benchmarking.
The optimizer uses
eggto explore equivalent query plans and selects a plan using the chosen cost model. Optimization is performed during query preparation; the selected plan is subsequently executed on the full graph by the existing LAGraph-based evaluator.Cost-based query optimization
The RPQMatrix implementation is reorganized into separate modules for query translation, rewrite rules, optimization, cost estimation, statistics, sampling, and execution.
The optimizer supports the following CLI strategies:
The rewrite rules cover reassociation of concatenation, distribution and factoring of alternatives, and replacement of expressions containing Kleene closure with specialized left- and right-closure operators. The rule set is initialized once and reused.
Cost models estimate both operation work and intermediate matrix characteristics. These estimates guide plan selection; they are not predictions of execution time in milliseconds.
MNC
The MNC model maintains row and column element counts, supplemented by singleton-related structural statistics. It uses these vectors to estimate multiplication and union cardinalities and propagates estimated statistics to parent expressions.
The implementation adapts MNC to Boolean RPQ relations rather than reproducing the original method in full. Intermediate counters may contain fractional estimates, and closure operators use fallback heuristics rather than a dedicated MNC closure model.
Pang-hybrid
The Pang-inspired model introduces an explicit estimate of closure growth instead of relying solely on a constant closure penalty. It simulates the growth of successive relation powers, estimates overlap with the accumulated result, and accounts for the work of closure steps.
The model uses scalar statistics rather than explicit vertex-domain intersections or Monte Carlo graph traversals. Its iteration limits apply only to cost estimation and do not restrict the paths considered during actual query execution.
Sampling
The sampling model selects a shared vertex subset for the labels used by a query and extracts the corresponding induced-subgraph matrices. Named query endpoints are explicitly included in the sample.
Boolean multiplication, union, and closure are performed on these smaller matrices during cost evaluation. The observed cardinalities and nonempty row and column counts are extrapolated to the full graph and combined with the analytical cost model.
Sampling is configurable through:
RPQ_SAMPLE_PERCENT: sampled vertex percentage, defaulting to 1%.RPQ_SAMPLE_SEED: reproducible random seed.RPQ_SAMPLE_MAX_STAR_ITERS: iteration limit for sampled closure evaluation.This is an approximate estimation technique: paths through excluded vertices are absent from the sample. Analytical estimates are retained when a sampled result is unavailable, an incomplete sample produces no matches, or closure evaluation has not converged. Sampling affects plan selection only; final answers are computed on the full graph.
Graph metadata and storage
The in-memory graph representation is extended with per-label matrix metadata:
The MatrixMarket loader computes and stores these statistics for reuse during optimization. Statistics collection is configurable through
MatrixStatsMode; the CLI requests extended statistics for MNC without imposing that additional work on other strategies.The optimizer can reuse precomputed statistics or construct them on demand when the graph backend does not provide them. Cached label statistics are associated with their underlying graph matrices to avoid reusing statistics from a different graph.
The MatrixMarket loader also maintains CSR and CSC matrix representations. Queries with a variable subject and a fixed object request CSC storage; other endpoint combinations request CSR. The graph interface provides a storage-aware lookup with a fallback for backends that do not maintain both orientations.
Fixed endpoints are represented by diagonal selector matrices and remain part of the expression processed by the optimizer.
Native GraphBLAS/LAGraph integration
The LAGraph dependency and Rust bindings are updated to support:
The native build script now tracks relevant LAGraph source, header, and CMake files so changes to the bundled dependency trigger a rebuild.
Benchmarking
This PR adds a fixed-run benchmark mode alongside Criterion-based benchmarking.
The fixed-run mode supports configurable warm-up and measured run counts and records two timing scopes separately:
These scopes are measured in separate loops, not as components of the same execution. FFI-only timing includes the execution wrapper and result handling performed by
execute, rather than isolating a single native call.Aggregate statistics are written to the main result file, while individual measured durations are saved in a companion
.runs.jsonfile. Benchmark metadata records the selected optimizer and measurement configuration. Checkpoint validation also includes the optimizer and benchmark settings to reject incompatible resume configurations.