mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-15 18:13:29 +02:00
* Update Q4_K and Q5_K to use branchless computation, which stops the scale unpack being re-executed for every column in mmvq, improving perf at batch sizes > 1 * Gating the change off from DGX Spark due to no gain * Adding prefetch gated to Spark, making branchless change in Q4_K and Q5_K general and modifying switch points based on latest perf data * Guard the mmvq L2 prefetch against MUSA as well as HIP * Define the mmvq L2 prefetch only under the Spark guard * Update switch point for Q4_K to accommodate more models * Remove stale comments * Add block_size to ggml_cuda_type_traits and create a separate mmvq_should_prefetch function * Rename block_size to bs for cleaner indentation * Fix build error on non-Spark CUDA arch with appropriate conditional around new function added --------- Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com>