Skip to content

GPU: skip events for default kernel launches - #9436

Draft
joseph-isaacs wants to merge 1 commit into
developfrom
joe/gpu-default-launch-no-events
Draft

GPU: skip events for default kernel launches#9436
joseph-isaacs wants to merge 1 commit into
developfrom
joe/gpu-default-launch-no-events

Conversation

@joseph-isaacs

@joseph-isaacs joseph-isaacs commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Independent of the other follow-up PRs; based directly on #9147. This draft contains one reviewable GPU correctness or performance concern.

What

  • Make kernel event recording conditional on the launch strategy.
  • Record no before/after events for normal execution.
  • Preserve timed events for tracing and profiling strategies.

Why this is needed

The default launch strategy records before/after CUDA events around every kernel even though its completion hook is a no-op. Those events therefore cannot affect correctness or produce a metric. The profile observed 1,865 event records across the benchmark. This change skips them only for the default strategy and preserves event timing for tracing/profiling strategies, removing two unnecessary driver operations per normal launch.

Validation

Checked independently against this PR's current base:

  • cargo check -p vortex-cuda --all-features

The combined implementation also passed Rust/CUDA formatting, 164 focused dynamic-dispatch tests, and 28 focused constant-array tests.

@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-nullable-dictionary-dispatch branch from e04e1fa to bf50ff8 Compare August 17, 2026 09:16
@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-default-launch-no-events branch from 179b26e to 67bb478 Compare August 17, 2026 09:17
@joseph-isaacs
joseph-isaacs changed the base branch from joe/gpu-nullable-dictionary-dispatch to claude/gpu-decompress-benchmarks-4mmn93 August 17, 2026 09:20
@codspeed-hq

codspeed-hq Bot commented Aug 17, 2026

Copy link
Copy Markdown

Merging this PR will degrade performance by 11.93%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

❌ 2 regressed benchmarks
✅ 2041 untouched benchmarks
⏩ 46 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
WallTime words_gather_scalar[65536] 8.2 µs 9.4 µs -12.74%
WallTime words_gather_dispatch[1024] 8 ns 9 ns -11.11%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing joe/gpu-default-launch-no-events (79352b6) with develop (b825c4f)

Open in CodSpeed

Footnotes

  1. 46 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports.

@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-default-launch-no-events branch from 67bb478 to bc6ea8d Compare August 17, 2026 12:45
@joseph-isaacs
joseph-isaacs changed the base branch from claude/gpu-decompress-benchmarks-4mmn93 to develop August 17, 2026 12:45
Signed-off-by: Joe Isaacs <2413449+joseph-isaacs@users.noreply.github.com>
@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-default-launch-no-events branch from bc6ea8d to 79352b6 Compare August 17, 2026 20:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant