Ollama
Tool, inference server, by Ollama
Assessment
The easiest way to serve a model on a desktop, especially on Windows. Use it for the helper tier and talk to its native API when a model has a thinking mode. Expect the one-model-at-a-time rule on a small card.
2026-10-09
Strengths
- One command to pull and serve a model on Windows, macOS or Linux
- A model library with sensible defaults, no GGUF hunting
- A native API that exposes model loading and thinking controls
Limitations
- Its OpenAI-compatible endpoint silently drops some fields, including the one that turns thinking off
- Only one model fits on a small card, and loading another evicts it
- The model name in a request must match a pulled tag
Ollama packages llama.cpp into a service with a model library: ollama pull qwen2.5-coder:3b, and a minute later there is an HTTP server answering chat requests.[1] It runs as a background service on Windows, macOS and Linux, keeps models loaded for a configurable time, and exposes both its own API and an OpenAI-compatible one.
What it is for
The helper tier on a desktop with a modest card, where the ease of install matters more than the last few percent of throughput. It is also what many desktop chat clients expect to find.
Things to know
- Two APIs, not one. The OpenAI-compatible shim covers the common fields and silently ignores others.
chat_template_kwargsis one it drops, so a thinking model keeps thinking; the native/api/chatwith"think": falseis the switch.[2] - The model name matters. Unlike a bare llama-server, Ollama resolves the
modelfield to a pulled tag and errors on an unknown one. - One model at a time on a small card. A 6 GB card holds either the text helper or the vision model; a request for the other evicts the first and the next call waits for a reload. Keep tags in config, send no
keep_aliveornum_ctxunless you mean to change what stays loaded, and never cause a load from a background job. OLLAMA_KEEP_ALIVE=-1on the server pins the loaded model for good; a per-request value overrides it for that model.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.