Qwen2.5 3B Instruct
Model, local llm, by Alibaba Qwen
Assessment
The right last-resort tier: a 3B instruct model at Q4 answers short structured calls at about 30 tokens a second on a desktop CPU and never needs the GPU. Keep it for classification, extraction and fallback, not for thinking.
2026-10-09
Strengths
- Short structured answers from a CPU with no GPU at all
- Four parallel slots in a few gigabytes of RAM
- Predictable output: not a reasoning model, so no hidden thinking to budget for
Limitations
- Anything that needs more than a paragraph of judgment
- Long context: 8K tokens is the practical ceiling on a CPU
Qwen2.5 3B Instruct is the smallest general-purpose model in Alibaba's Qwen2.5 family that still follows instructions reliably.[1] At a 4-bit quantization the weights fit in about two gigabytes, which is what makes it useful as a tier that lives entirely on a CPU and stays resident while a GPU is lent to something else.
What it is for
Short, structured work: classify a message, extract fields into JSON under a grammar, write a one-line summary, answer a yes-or-no question about a document. It is not a reasoning model, so there is no hidden chain of thought to turn off and no budget surprise; the first tokens it writes are the answer.
How it runs
Through llama-server with no GPU layers, 8K context per slot and four slots, it decodes at roughly 30 tokens a second when idle and about 10 when all four slots are busy.[2] A CPU instance survives any GPU swap, which is the whole reason to keep one: it is the tier that answers when nothing else can.
Things to know
- The server ignores the model name in a request; send an honest one for the logs anyway.
- Quality drops quickly past a few hundred tokens of output. Ask it small questions.
- The Coder variant of the same size is the better pick when the work is code or JSON.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.