Skip to content

Optimize channelize_poly CUDA kernels and dispatch - #1253

Merged
tbensonatl merged 1 commit into
mainfrom
tbenson/channelize-poly-fft-fusion-opts
Oct 2, 2026
Merged

tbensonatl merged 1 commit into
mainfrom
tbenson/channelize-poly-fft-fusion-opts

Conversation

@tbensonatl

Copy link
Copy Markdown
Collaborator

Add fused FIR+DFT kernels that keep filtered samples in shared memory or registers, avoiding the intermediate tensor and separate cuFFT launch. Add small-channel leaves for critically sampled M=2..6 (replacing the prior FusedChan kernel) and general leaves for M=8, 10, 16, 20, 32, 40, 64, and 80, plus oversampled M=3..6. The infrastructure is in-place to extend coverage of fusion, but this PR does not do so in order to limit template instantiations. Instead, the aim is for a future PR to add a compile-time property that allows users to opt-in specific channel counts for fusion.

Rework the FIR+cuFFT backends via tuning of tile sizes and launch configuration. Select backends through SelectPlan/ExecutePlan using device attributes (SM count, L2 size, FP64 throughput, memory bus width).

Tests compare every launchable backend, including windowed launches, against the host implementation.

Across a 4,832-shape sweep, the geometric-mean speedup over the previous implementation is 1.89x on L4 and 2.15x on GH200. Some regressions remain with these changes, but the vast majority of cases improve with some cases being more than twice as fast.

Add fused FIR+DFT kernels that keep filtered samples on chip, avoiding the intermediate tensor
and separate cuFFT launch: small-channel leaves for critically sampled M=2..6 (replacing
FusedChan) and general leaves for M=8, 10, 16, 20, 32, 40, 64, and 80, plus oversampled M=3..6.

Rework the FIR+cuFFT backends: tuning of tile sizes and launch
configuration.

Select backends through SelectPlan/ExecutePlan using device attributes (SM count, L2 size, FP64
throughput, memory bus width).

Tests compare every launchable backend, including windowed launches, against the host
implementation. channelize_poly_bench gains single-case options.

Across a 4,832-shape sweep, the geometric-mean speedup over the previous implementation is
1.89x on L4 and 2.15x on GH200.

Signed-off-by: Thomas Benson <tbenson@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Optimizes polyphase channelizer kernels and updates tests.

The PR appears safe to merge; no actionable regression was established.

Summary

The PR adds fused CUDA FIR+DFT channelizer kernels and selects among fused and separate backends using shared-memory fit and device characteristics. It also retunes FIR launches, extends backend and windowed-output tests, and documents half-precision limitations.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Validate channelizer arguments] --> B{Fused leaf feasible and preferred?}
  B -- Yes --> C[Fused FIR and DFT]
  B -- No --> D{Critical sampling and Smem suitable?}
  D -- Yes --> E[Whole-channel shared-memory FIR]
  D -- No --> F{Tiled plan fits?}
  F -- Yes --> G[Channel-tiled FIR]
  F -- No --> H[Generic FIR]
  E --> I[cuFFT]
  G --> I
  H --> I
  I --> J{Real filtered samples?}
  J -- Yes --> K[Unpack spectrum]
  J -- No --> L[Output]
  K --> L
  C --> L
Loading

Reviews (1) · Last reviewed commit: "Optimize channelize_poly CUDA kernels an..."

@tbensonatl

Copy link
Copy Markdown
Collaborator Author

/build

@tbensonatl tbensonatl self-assigned this Oct 2, 2026
@coveralls

Copy link
Copy Markdown

Coverage Status

Coverage is 93.629% — tbenson/channelize-poly-fft-fusion-opts into main. No base build found for main.

@tbensonatl
tbensonatl merged commit ce764c0 into main Oct 2, 2026
2 checks passed
@tbensonatl
tbensonatl deleted the tbenson/channelize-poly-fft-fusion-opts branch October 2, 2026 16:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants