mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-01 01:27:50 +02:00
* common : dedupe --n-cpu-moe / --spec-draft-n-cpu-moe override loops * common : add --n-cpu-ffn to CPU-offload dense FFN weights of first N layers * common : generalize llm_ffn_block_regex over the FFN regex, drop TODO