> vectormatrix.wiki

Ollama

Tool, inference server, by Ollama

Assessment

The easiest way to serve a model on a desktop, especially on Windows. Use it for the helper tier and talk to its native API when a model has a thinking mode. Expect the one-model-at-a-time rule on a small card.

2026-10-09

Strengths

  • One command to pull and serve a model on Windows, macOS or Linux
  • A model library with sensible defaults, no GGUF hunting
  • A native API that exposes model loading and thinking controls

Limitations

  • Its OpenAI-compatible endpoint silently drops some fields, including the one that turns thinking off
  • Only one model fits on a small card, and loading another evicts it
  • The model name in a request must match a pulled tag

Ollama packages llama.cpp into a service with a model library: ollama pull qwen2.5-coder:3b, and a minute later there is an HTTP server answering chat requests.[1] It runs as a background service on Windows, macOS and Linux, keeps models loaded for a configurable time, and exposes both its own API and an OpenAI-compatible one.

What it is for

The helper tier on a desktop with a modest card, where the ease of install matters more than the last few percent of throughput. It is also what many desktop chat clients expect to find.

Things to know

  • Two APIs, not one. The OpenAI-compatible shim covers the common fields and silently ignores others. chat_template_kwargs is one it drops, so a thinking model keeps thinking; the native /api/chat with "think": false is the switch.[2]
  • The model name matters. Unlike a bare llama-server, Ollama resolves the model field to a pulled tag and errors on an unknown one.
  • One model at a time on a small card. A 6 GB card holds either the text helper or the vision model; a request for the other evicts the first and the next call waits for a reload. Keep tags in config, send no keep_alive or num_ctx unless you mean to change what stays loaded, and never cause a load from a background job.
  • OLLAMA_KEEP_ALIVE=-1 on the server pins the loaded model for good; a per-request value overrides it for that model.

History

  • 2026-10-09: article written and published.
  • 2026-10-09: entry created.

Practices

References

  1. Ollama ^
  2. Ollama API reference ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.