llama.cpp and llama-server
Tool, inference server, by ggerganov and contributors
Assessment
The inference server for everything local here, GPU and CPU alike. Native llama-server replaced the Python wrapper in July 2026 and made grammar output nearly free. Watch the slot count and health-check the model rather than the process.
2026-10-09
Strengths
- Running GGUF models on consumer GPUs and CPUs with one binary
- An OpenAI-compatible HTTP server with grammar-constrained JSON output
- Multi-GPU tensor splitting and a resident vision adapter
Limitations
- Its slots are the whole capacity; the default one slot serializes every caller
- The model field is ignored, so a wrong model name never errors
llama.cpp is the C/C++ inference engine behind most local LLM setups,[1] and llama-server is its HTTP front: OpenAI-compatible chat completions, embeddings, /health and /slots endpoints, JSON-schema output through GBNF grammars, and vision through a separate mmproj file.[2]
Why it, and not a wrapper
The same model served through llama-cpp-python's server paid a flat 3.7× per-token cost under a grammar, because logit masking over a 256K vocabulary was done in Python. Native llama-server does it in C, so constrained and free text run at the same speed. The two servers also take different response_format shapes, and each silently ignores the other's, which is the kind of thing worth keeping in an environment variable.
Things to know
--parallelis the number of slots, and the default is one. CheckGET /slotsbefore putting an interactive user behind a batch job.- A stopped unit does not always refuse connections; set a short connect timeout.
- Expect a minute or two of 503s after a restart while the model loads. That is not an outage.
- A CPU-only instance (
--n-gpu-layers 0) of a 3B model is a usable last-resort tier that survives the GPU being lent to something else.
History
- 2026-10-09: record seeded.
- 2026-07-06: native llama-server replaced llama-cpp-python's server on the primary box.
Practices
- Turn the reasoning channel off on Gemma 4, in every request
- Catch the transport error family and set a short connect timeout
- A JSON schema constrains the answer, not the thinking
- Pick the host before you stream, and prewarm
- One inference slot is one caller for the whole system
- Match response_format to the server that is actually running
- Health-check the model, not the process
- Handle the empty 200, not just the connection error
- Grammar-constrained output is nearly free on native llama-server
Comments
Public comments on each entry are coming. Nothing is collected here yet.