GPU: keep patch indices on device - #9433
Performance Regression: -14.03%
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
❌ 7 regressed benchmarks
✅ 2036 untouched benchmarks
⏩ 46 skipped benchmarks1
Warning
Please fix the performance issues or acknowledge them on CodSpeed.
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | cuda/alp_f64/10%[100M] |
6.9 ms | 8.4 ms | -17.79% |
| ❌ | WallTime | cuda/alp_f32/0%[100M] |
2.5 ms | 2.9 ms | -14.98% |
| ❌ | WallTime | cuda/alp_f64/1%[100M] |
6.6 ms | 7.7 ms | -14.36% |
| ❌ | WallTime | cuda/alp_f32/1%[100M] |
4.8 ms | 5.6 ms | -14.28% |
| ❌ | WallTime | cuda/alp_f32/10%[100M] |
4.5 ms | 5.1 ms | -13.01% |
| ❌ | WallTime | words_gather_scalar[65536] |
8.2 µs | 9.4 µs | -12.49% |
| ❌ | WallTime | words_gather_dispatch[1024] |
8 ns | 9 ns | -11.11% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing joe/gpu-device-patch-indices (67c9110) with develop (b825c4f)
Footnotes
-
46 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩