> vectormatrix.wiki

llama.cpp and llama-server

Tool, inference server, by ggerganov and contributors

Assessment

The inference server for everything local here, GPU and CPU alike. Native llama-server replaced the Python wrapper in July 2026 and made grammar output nearly free. Watch the slot count and health-check the model rather than the process.

2026-10-09

Strengths

  • Running GGUF models on consumer GPUs and CPUs with one binary
  • An OpenAI-compatible HTTP server with grammar-constrained JSON output
  • Multi-GPU tensor splitting and a resident vision adapter

Limitations

  • Its slots are the whole capacity; the default one slot serializes every caller
  • The model field is ignored, so a wrong model name never errors

llama.cpp is the C/C++ inference engine behind most local LLM setups,[1] and llama-server is its HTTP front: OpenAI-compatible chat completions, embeddings, /health and /slots endpoints, JSON-schema output through GBNF grammars, and vision through a separate mmproj file.[2]

Why it, and not a wrapper

The same model served through llama-cpp-python's server paid a flat 3.7× per-token cost under a grammar, because logit masking over a 256K vocabulary was done in Python. Native llama-server does it in C, so constrained and free text run at the same speed. The two servers also take different response_format shapes, and each silently ignores the other's, which is the kind of thing worth keeping in an environment variable.

Things to know

  • --parallel is the number of slots, and the default is one. Check GET /slots before putting an interactive user behind a batch job.
  • A stopped unit does not always refuse connections; set a short connect timeout.
  • Expect a minute or two of 503s after a restart while the model loads. That is not an outage.
  • A CPU-only instance (--n-gpu-layers 0) of a 3B model is a usable last-resort tier that survives the GPU being lent to something else.

History

  • 2026-10-09: record seeded.
  • 2026-07-06: native llama-server replaced llama-cpp-python's server on the primary box.

Practices

References

  1. llama.cpp on GitHub ^
  2. llama-server README ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.