ggml-cpu: dequantize each weight row once per matmul tile (4-column Q4_0/Q4_K AVX2 kernels) - #41
Open
codspeed-hq[bot] wants to merge 1 commit into
CodSpeed HQ / CodSpeed Performance Analysis
succeeded
Jul 29, 2026 in 0s
Performance Gate Passed
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 2 improved benchmarks
✅ 26 untouched benchmarks
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | Simulation | mul_mat[q4_0] |
102 ms | 59.6 ms | +71.03% |
| ⚡ | Simulation | mul_mat[q4_k] |
72.4 ms | 46.7 ms | +54.91% |
Tip
Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.
Comparing codspeed-optim-dequantize-each-weight-row-once-per-matmul-tile-4-1785338847703 (7baa65d) with master (46819c9)
Loading