A `setParams` / `setParamsByID` key ending in `?` is set-if-undefined: the value is applied only when the request doesn't already carry that parameter. Plain keys keep forcing their values, and a config that doesn't use the suffix behaves exactly as before. Fixes: #1052
Hello!
Here you will find the knowledge base for llama-swap. It's easy to get started in llama-swap with just a few lines of YAML. However, the real power comes from the dozens of configuration options to control routing and resource loading exactly as you want it.
llama-swap doesn't come with traditional documentation. Instead, it includes a documentation agent that reads from the knowledge base to answer your questions directly.
Three steps to get started:
- Download gemma-4-12B
- Install llama-server
- Write your first configuration file and start llama-swap
Downloading gemma-4-12B
Docs is evaluated against a gemma-4-12B Q4_K_M from Unsloth. It is a small and capable model.
Download it form https://huggingface.co/unsloth/gemma-4-12b-it-GGUF
uvx hf download unsloth/gemma-4-12b-it-GGUF gemma-4-12b-it-Q4_K_M.gguf --local-dir .
# download MTP
uvx hf download unsloth/gemma-4-12b-it-GGUF MTP/mtp-gemma-4-12b-it-Q8_0.gguf --local-dir .
If you have a smaller GPU or none at all, use the lighter gemma-4-E4B instead.
Installing llama-server
(find instructions for your os) - to be written.
Installing llama-swap
(to be written)
config.yaml
Use this minimal configuration to get gemma-4-12B running to power the Docs agent.
models:
gemma-4-12B:
cmd: |
/path/to/llama-server-latest
--host 127.0.0.1 --port ${PORT}
--log-verbosity 4 --log-colors on
--temp 1.0 --top-p 0.95 --top-k 64
# model params
--model /path/to/gemma-4-12b-it-Q4_K_M.gguf
--model-draft /path/to/mtp-gemma-4-12b-it-Q8_0.gguf
# enable MTP
--spec-type draft-mtp
--spec-draft-n-max 4 --spec-draft-p-min 0.75
capabilities:
tools: true
filters:
stripParams: "temperature, top_k, top_p, repeat_penalty, min_p, presence_penalty"
Run llama-swap
Start up llama-swap and visit http://localhost:8080
llama-swap -config config.yaml -listen localhost:8080