Add experimental sparse matrix-matrix multiplication - #1215
liaoweiyang2017 wants to merge 7 commits into
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #1215 +/- ##
==========================================
- Coverage 68.20% 68.12% -0.09%
==========================================
Files 19 19
Lines 2378 2378
==========================================
- Hits 1622 1620 -2
- Misses 756 758 +2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
Thanks for the proposal @liaoweiyang2017 Just a few first comments before a full review:
Why? For performance sake, it might be desired to call the In other scenarios, like computing the exponent of a matrix, indeed an incremental increase of the matrix sparsity pattern might be desired, so we need a way to control whether to just let target matrix to be resized or just recycle the same existing matrix. I see you managed the combinations of different sparse types, I would have kept the PR reduced to same-type products, open for debate. |
|
Thanks @jalvesz for the suggestions. I have pushed an update in commit
The symbolic pattern includes stored zeros and positions whose numerical contributions cancel, so changes in numerical values do not force preparation again. A change to indices or storage layout requires a new preparation; dimensions alone are not sufficient to establish compatibility. I added tests for numerical reuse, unchanged storage addresses, structure mismatches, explicit rebuilding, zero scaling, cancellation, invalid storage and empty products, and updated the examples and specification. The local sparse suite passes in FPM Debug and CMake Release, including the configured real/complex kinds. An allocator-counting probe also records zero allocations for repeated prepared calls in each of the five formats. A complete 50,000-by-50,000 CSC result agrees with Julia in both pattern and values. The reusable path reduces repeated-call time on the local fixture. Preparation adds time and integer memory compared with the previous fused one-shot implementation; I will keep that distinction explicit in the validation notes. A separate comparison now covers six large CSC fixtures. The Fortran prepared path is close to a local Julia reference implementing the same checked numerical-reuse workflow, with about 1.4%–13.0% higher call times in these measurements. That Julia reference is not an official SparseArrays API. Comparisons with the official allocating The names and plan-based interface are still experimental, and I would appreciate feedback on this separation before considering the PR ready for a full review. |
Declare shared helpers as private separate module procedures and implement them in a common submodule. GCC 13/14 otherwise localize or remove contained private helpers referenced by format submodules. Remove the duplicate PUBLIC attribute on spmm_plan_type rejected by Intel 2024.1.
The Windows OpenBLAS failure is a relative-only comparison near zero. Add a deterministic cancellation regression and a precision/data-scale absolute error bound to the scaled symmetric tridiagonal comparisons. Keep the production multiplication code unchanged.
Use three distinct keys with a fixed colliding hash to exercise relocation after removal from a wrapped probe run. Assert key absence, survivor reachability, entry count and stored values. This stabilizes the two-line coverage variation reported by codecov/project without changing production code or coverage thresholds.
|
A first general comment: for the sake of transparency, could you please mention which Gen AI system are you using to help you with the development? Coming back to the subject. I have the impression that the current design uses OOP quite heavily and I wonder if it is truly needed? Some aspects do require but I believe that OOP should be kept at minimum, specially in performance sensitive areas of code where plain array manipulations could be enough. I'm concerned about why the split of matrix creation and application of the actual kernel introduced a measurable loss in performance. I would have imagined:
Before modifying your current proposal, could you evaluate this comment and give a feedback on what it would imply in terms of both performance and code quality? Thanks! |
Use caller-owned result structure and optional integer workspace instead of persistent structural snapshots. Add independent N/T/H operations, plain-array kernels for all five formats, and checks/tests documenting the caller-owned pattern contract. Keep this candidate local pending the requested design discussion.
|
Thanks @jalvesz for the feedback. I am using OpenAI Codex to assist with the design, implementation, tests and benchmark harnesses. The reported results come from actual local compiler/test runs, and I have added this disclosure to the PR description. I evaluated the array-based design locally and have now pushed the tested candidate in I agree that the persistent plan was more machinery than this interface needs. The previous numerical path was not doing virtual dispatch for every multiplication; the main removable work was structural snapshots, full index comparisons on every call, identity slot maps for CSR/CSC, and repeated copying during preparation. The revision has no custom plan/structure types or polymorphic dispatch in the SpMM implementation:
The contract is intentionally lighter: the high-level wrapper checks operations, dimensions, storage and buffer/workspace extents. The raw kernel assumes conforming storage and a C pattern containing every required output position. A valid precomputed superset is accepted. Changes to values alone need no preparation; changes to indices or operations require the caller to establish that C is still sufficient or call preparation again. Exact structural-change detection is no longer provided automatically. On the same Apple M5 Pro with GNU Fortran 13.4.0,
The original fused implementation, rebuilt with the same compiler/options, took 92.9 ms for the complete call. Thus removing the extra state recovers substantial performance, but a separate symbolic traversal and initialization still have a measurable one-shot cost. Iterative use amortizes preparation. Preserving the best fused one-shot time as well would require an additional execution path and its maintenance cost. For the same large case, Julia's official allocating multiplication took 160.0 ms. A separately labelled local Julia array reference took 44.6 ms for the numerical kernel; this is not an official SparseArrays API. At 50,000 square and 10 entries per column, the Fortran complete call took 13.2 ms versus Julia's 11.2 ms. These are results for one structured family on one machine, not a general speed ranking. The direct N×N CSR/CSC/ELL/SELLC paths allocate nothing when workspace is supplied. COO grouping and transposed inputs still use transient integer views; they do not reallocate C or copy input numerical values. A strict allocation-free requirement for every format/operation would need an explicit larger workspace layout or stronger input-ordering requirements. For the large CSC example, the old plan's integer storage was about 302 MiB by array-size accounting; the direct numeric workspace is about 0.19 MiB, excluding the matrices and preparation temporaries. Locally, the 17 sparse tests pass with GNU 13/16 runtime checking and GNU 13 FPM Release (GNU 13 includes real/complex SP/DP/QP; the GNU 16 build covers SP/DP). They cover all five formats, all N/T/H pairs, rectangular and empty cases, and result storage reuse. Both examples and the FORD build pass. Six large cases produced 36 complete pointer/index/value comparisons against the previous implementation and Julia, with no mismatches at In code-quality terms, the persistent state and full-structure validation disappear, while responsibility for a reusable result pattern becomes explicit. The source is slightly longer because it adds independent operations and five public array interfaces; the high-level usage becomes simpler. I would appreciate your view on this contract and on whether the initial implementation should prioritize this two-phase reuse path or also retain a fused one-shot specialization. |
|
Thanks for the detailed explanation and figures. I guess that, given the numbers, we could reconsider coming back to a fused call and keep the boolean flag for avoiding the reallocations if needed. That way the C matrix would be formed once already with the first application and subsequent calls can be more economical by passing the boolean to .false. What do you think about it @liaoweiyang2017 ? Maybe @jvdp1 has a different take on this? |
Summary
This draft adds experimental
spmmsupport following the SpMV layout: one public module, five format-specific submodules, and a shared private implementation submodule. The array-kernel revision is in53a9bb70;3d1e5158only relocates the changelog entry to avoid an upstream conflict.sparse_fullstorage.spmm_preparebuilds C independently of numerical values, with independent N/T/H operations on A and B.spmm_kernel_*interfaces operate on plain arrays and an integer workspace. No persistent plan or structural snapshots are used.spmmcalls preparation by default;allow_resize=.false.reuses the supplied result structure.The caller must ensure that C contains the complete product pattern; a valid superset is accepted. Structural changes are not detected automatically. The high-level wrapper checks operations, dimensions and storage/workspace extents; raw kernels assume conforming arrays. The result must not alias an input, and concurrent calls require separate result/work buffers. Stored zeros and cancellation do not remove prepared positions.
Local validation
rtol=atol=1e-12; maximum absolute difference4.44e-16.Performance evaluation
Apple M5 Pro, GNU Fortran 13.4.0
-O3, Float64/Int32 CSC, 5 warm-ups, 15 samples and three independent processes. At 50,000 square with 30 entries per input column:For the same fixture, the original fused Fortran implementation takes 92.9 ms for a complete call. Julia's official allocating multiplication takes 160.0 ms; a separate local Julia array reference takes 44.6 ms for numerical reuse (not an official SparseArrays API). At 50,000 square and 10 entries per column, the new Fortran complete call takes 13.2 ms versus Julia's 11.2 ms. These measurements cover one structured family on one machine, not a general ranking. Benchmark and allocator probes remain outside the PR diff.
The two-phase path improves repeated use but retains extra one-shot symbolic traversal and initialization. Whether to maintain a separate fused one-shot path remains open for discussion, as do the caller-owned pattern contract and same-format scope. Hosted CI for this revision requires maintainer approval; local passes do not establish hosted CI success.
AI assistance
OpenAI Codex assisted with design, implementation, tests, benchmark harnesses and documentation. The validation and timing results above come from actual local runs.