Skip to content

ggml-cpu: unpack the Q4_K scales/mins with SIMD in the AVX2 ggml_vec_dot_q4_K_q8_K - #35

Open
codspeed-hq[bot] wants to merge 1 commit into
masterfrom
codspeed-optim-unpack-the-q4-k-scales-mins-with-simd-in-avx2-ggml-1785298752860
Open

ggml-cpu: unpack the Q4_K scales/mins with SIMD in the AVX2 ggml_vec_dot_q4_K_q8_K#35
codspeed-hq[bot] wants to merge 1 commit into
masterfrom
codspeed-optim-unpack-the-q4-k-scales-mins-with-simd-in-avx2-ggml-1785298752860

ggml-cpu: unpack Q4_K scales/mins with SIMD in AVX2 ggml_vec_dot_q4_K…

3db79c4
Select commit
Loading
Failed to load commit list.
CodSpeed HQ / CodSpeed Performance Analysis succeeded Jul 29, 2026 in 0s

Performance Gate Passed

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 1 improved benchmark
✅ 27 untouched benchmarks

Performance Changes

Mode Benchmark BASE HEAD Efficiency
Simulation mul_mat[q4_k] 72.4 ms 63 ms +14.85%

Tip

Curious why this is faster? Comment @codspeedbot explain why this is faster on this PR, or directly use the CodSpeed MCP with your agent.


Comparing codspeed-optim-unpack-the-q4-k-scales-mins-with-simd-in-avx2-ggml-1785298752860 (3db79c4) with master (46819c9)

Open in CodSpeed