Qwen2.5 Coder 3B
Model, local llm, by Alibaba Qwen
Assessment
The best small helper for a 6 GB card: fast, loads in a few gigabytes, and strong at the structured, code-shaped work an agent fans out. Keep it resident and send it small questions.
2026-10-09
Strengths
- Code completion, small edits and JSON at about 48 tokens a second on a 6 GB card
- Stays loaded at a 32K context in about 3.4 GB of VRAM
- Fast fan-out: many small calls in parallel from an agent
Limitations
- Reasoning across a whole repository
- Shares a small card with any vision model, so a photo evicts it
Qwen2.5 Coder 3B is the code-specialised member of the Qwen2.5 family at the three-billion size.[1] It was trained on a large share of source code and code-adjacent text, which shows in two ways that matter for agents: it writes syntactically valid code and JSON more reliably than a general model of the same size, and it follows terse, tool-like instructions well.
What it is for
The helper role: the many small calls an agent makes around the main model. Rename a function, write a commit message, turn a sentence into a JSON object, pick a category, draft a prompt. Served through Ollama it decodes at about 48 tokens a second on a GTX 1060 and loads in about 3.4 GB at a 32K context, so it can stay resident on a card that size.[2]
Things to know
- A 6 GB card holds one model at a time. If a vision model is loaded for a photo, this one is evicted and the next call pays a reload of about twelve seconds.
- Send no
keep_aliveand nonum_ctxfrom a client unless you mean to change what stays loaded; the server pins the model and a different context length is a reload. - It is not a reasoning model and does not pretend to be. For judgment, go up a tier.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.