> vectormatrix.wiki

Gemma 4 26B A4B

Model, local llm, by Google

Assessment

The primary model on this system and the right one for a two-card home box. Its mixture-of-experts shape (26B total, about 4B active) gives 65 tokens a second at Q4 with 24K context. Turn thinking off per request and it answers short structured calls in about a second.

2026-10-09

Strengths

  • Long-context reasoning on a pair of consumer cards (24K context on 2× 12 GB)
  • Structured JSON output under a grammar at full speed on native llama-server
  • Reading photos through its vision adapter (OCR, receipts, odometers)

Limitations

  • Hidden chain-of-thought by default, which eats the token budget unless turned off
  • One inference slot per server, so a batch job stalls every interactive caller

Gemma 4 26B A4B is Google's open-weights mixture-of-experts model: 26 billion parameters in total, about four billion active per token, which is why it runs at interactive speed on two mid-range cards.[2] The instruction-tuned build is a reasoning model: it writes a hidden chain of thought before every answer unless told not to.

How it is run here

Native llama-server from llama.cpp, the Unsloth UD-Q4_K_M quantization,[1] split evenly across two 12 GB cards with flash attention on, 24,576 tokens of context and a single inference slot. The vision adapter (mmproj) is resident, so an OpenAI-style image_url content part works. Thinking is off server-side by default and every client still sends chat_template_kwargs: {"enable_thinking": false} per request, because a server swap would otherwise bring it back.

Measured

  • 65 tokens a second, free text and grammar-constrained alike, on native llama-server.
  • A grammar-constrained reminder call went from 600 tokens and 8.7 s with thinking on to 72 tokens and 1.4 s with it off.
  • A photo adds about 1.5 GB on one card and leaves decode speed unchanged.

History

  • 2026-10-09: record seeded from the system's LLM reference.
  • 2026-09-24: vision adapter made resident; the first image aborted the server until the micro-batch was raised to 512.
  • 2026-08-25: thinking turned off server-side; the quantization line corrected to UD-Q4_K_M.
  • 2026-07-06: moved from llama-cpp-python to native llama-server, which removed a 3.7× grammar tax.

Practices

References

  1. Unsloth GGUF (UD-Q4_K_M) on Hugging Face ^
  2. Gemma on Google AI for Developers ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.