ggml-cpu: unroll the AVX2 q4_K x q8_K sub-block loop - #44
Open
codspeed-hq[bot] wants to merge 1 commit into
Open
CodSpeed HQ / CodSpeed Performance Analysis
succeeded
Jul 30, 2026 in 0s
Performance Gate Passed
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 1 improved benchmark
✅ 27 untouched benchmarks
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | Simulation | mul_mat[q4_k] |
72.4 ms | 64.4 ms | +12.44% |
Tip
Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.
Comparing codspeed-optim-unroll-the-avx2-ggml-vec-dot-q4-k-q8-k-sub-block-s-1785403756979 (44a67ae) with master (46819c9)
Loading