GPU: skip events for default kernel launches - #9436
Conversation
e04e1fa to
bf50ff8
Compare
179b26e to
67bb478
Compare
Merging this PR will degrade performance by 11.93%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | words_gather_scalar[65536] |
8.2 µs | 9.4 µs | -12.74% |
| ❌ | WallTime | words_gather_dispatch[1024] |
8 ns | 9 ns | -11.11% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing joe/gpu-default-launch-no-events (79352b6) with develop (b825c4f)
Footnotes
-
46 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
67bb478 to
bc6ea8d
Compare
Signed-off-by: Joe Isaacs <2413449+joseph-isaacs@users.noreply.github.com>
bc6ea8d to
79352b6
Compare
Independent of the other follow-up PRs; based directly on #9147. This draft contains one reviewable GPU correctness or performance concern.
What
Why this is needed
The default launch strategy records before/after CUDA events around every kernel even though its completion hook is a no-op. Those events therefore cannot affect correctness or produce a metric. The profile observed 1,865 event records across the benchmark. This change skips them only for the default strategy and preserves event timing for tracing/profiling strategies, removing two unnecessary driver operations per normal launch.
Validation
Checked independently against this PR's current base:
cargo check -p vortex-cuda --all-featuresThe combined implementation also passed Rust/CUDA formatting, 164 focused dynamic-dispatch tests, and 28 focused constant-array tests.