ggml-quants: vectorize the per-block extremes scan in the reference quantizers - #43
Open
codspeed-hq[bot] wants to merge 1 commit into
CodSpeed HQ / CodSpeed Performance Analysis
succeeded
Jul 30, 2026 in 0s
Performance Gate Passed
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 4 improved benchmarks
✅ 24 untouched benchmarks
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | Simulation | quantize_chunk[q4_0] |
246.7 µs | 140.2 µs | +75.95% |
| ⚡ | Simulation | quantize_chunk[q4_1] |
213.6 µs | 144.8 µs | +47.53% |
| ⚡ | Simulation | quantize_chunk[q5_0] |
386 µs | 279.2 µs | +38.25% |
| ⚡ | Simulation | quantize_chunk[q8_0] |
175 µs | 135.6 µs | +29.1% |
Tip
Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.
Comparing codspeed-optim-vectorize-the-per-block-extremes-scan-in-the-refer-1785394478758 (15ee065) with master (46819c9)
Loading