Skip to content

GPU: materialize flat constant arrays on device - #9434

Draft
joseph-isaacs wants to merge 1 commit into
developfrom
joe/gpu-flat-constant-arrays
Draft

GPU: materialize flat constant arrays on device#9434
joseph-isaacs wants to merge 1 commit into
developfrom
joe/gpu-flat-constant-arrays

Conversation

@joseph-isaacs

@joseph-isaacs joseph-isaacs commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Independent of the other follow-up PRs; based directly on #9147. This draft contains one reviewable GPU correctness or performance concern.

What

  • Materialize null, boolean, primitive, decimal, UTF-8, binary, and flat extension constants as device-resident canonical arrays.
  • Preserve nullable validity, including all-invalid constants.
  • Add focused coverage for inline and outlined values.

Why this is needed

Constant nodes are normal compressor output: constant columns, all-null columns, and constant patch values all produce ConstantArray. The base CUDA executor only handles non-null numeric/decimal constants, so Boolean, UTF-8, binary, extension, and null constants fall back to CPU inside an otherwise GPU-resident tree. The profile exposed a concrete “Nullable UTF-8 Constant CPU fallback” costing 14.573 ms, making this both a correctness-coverage and measured performance fix.

Validation

Checked independently against this PR's current base:

  • cargo check -p vortex-cuda --all-features

The combined implementation also passed Rust/CUDA formatting, 164 focused dynamic-dispatch tests, and 28 focused constant-array tests.

@codspeed-hq

codspeed-hq Bot commented Aug 17, 2026

Copy link
Copy Markdown

Merging this PR will degrade performance by 11.78%

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

❌ 2 regressed benchmarks
✅ 2041 untouched benchmarks
⏩ 46 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
WallTime words_gather_scalar[65536] 8.2 µs 9.4 µs -12.45%
WallTime words_gather_dispatch[1024] 8 ns 9 ns -11.11%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing joe/gpu-flat-constant-arrays (ee5e383) with develop (b825c4f)

Open in CodSpeed

Footnotes

  1. 46 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports.

@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-device-patch-indices branch from 171fc0f to 32e03e6 Compare August 17, 2026 09:15
@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-flat-constant-arrays branch from e271be1 to 1338a6a Compare August 17, 2026 09:15
@joseph-isaacs
joseph-isaacs changed the base branch from joe/gpu-device-patch-indices to claude/gpu-decompress-benchmarks-4mmn93 August 17, 2026 09:20
@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-flat-constant-arrays branch from 1338a6a to 559fa98 Compare August 17, 2026 12:42
@joseph-isaacs
joseph-isaacs changed the base branch from claude/gpu-decompress-benchmarks-4mmn93 to develop August 17, 2026 12:42
@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-flat-constant-arrays branch from 559fa98 to 4bbfa3c Compare August 17, 2026 18:31
Signed-off-by: Joe Isaacs <2413449+joseph-isaacs@users.noreply.github.com>
@joseph-isaacs
joseph-isaacs force-pushed the joe/gpu-flat-constant-arrays branch from 4bbfa3c to ee5e383 Compare August 17, 2026 19:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant