> vectormatrix.wiki

Qwen2.5 3B Instruct

Model, local llm, by Alibaba Qwen

Assessment

The right last-resort tier: a 3B instruct model at Q4 answers short structured calls at about 30 tokens a second on a desktop CPU and never needs the GPU. Keep it for classification, extraction and fallback, not for thinking.

2026-10-09

Strengths

  • Short structured answers from a CPU with no GPU at all
  • Four parallel slots in a few gigabytes of RAM
  • Predictable output: not a reasoning model, so no hidden thinking to budget for

Limitations

  • Anything that needs more than a paragraph of judgment
  • Long context: 8K tokens is the practical ceiling on a CPU

Qwen2.5 3B Instruct is the smallest general-purpose model in Alibaba's Qwen2.5 family that still follows instructions reliably.[1] At a 4-bit quantization the weights fit in about two gigabytes, which is what makes it useful as a tier that lives entirely on a CPU and stays resident while a GPU is lent to something else.

What it is for

Short, structured work: classify a message, extract fields into JSON under a grammar, write a one-line summary, answer a yes-or-no question about a document. It is not a reasoning model, so there is no hidden chain of thought to turn off and no budget surprise; the first tokens it writes are the answer.

How it runs

Through llama-server with no GPU layers, 8K context per slot and four slots, it decodes at roughly 30 tokens a second when idle and about 10 when all four slots are busy.[2] A CPU instance survives any GPU swap, which is the whole reason to keep one: it is the tier that answers when nothing else can.

Things to know

  • The server ignores the model name in a request; send an honest one for the logs anyway.
  • Quality drops quickly past a few hundred tokens of output. Ask it small questions.
  • The Coder variant of the same size is the better pick when the work is code or JSON.

History

  • 2026-10-09: article written and published.
  • 2026-10-09: entry created.

References

  1. Qwen2.5-3B-Instruct GGUF ^
  2. Qwen2.5 technical report ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.