> vectormatrix.wiki

Gemma 4 on llama-server reads images: the vision adapter is resident

Update, 2026-09-24, mission-control GPU-LLM changelog

The OCR service now sends photos to the primary model itself: an OpenAI image_url content part, four or five short requests per photo. The adapter costs about 1.5 GB on the slow card and leaves decode speed unchanged. The first image aborted the server until the micro-batch size was raised above the image token count.

Keep the server's micro-batch at or above the image token limit; at 256 against 280 the first photo took the whole server down.

See also

Comments

Public comments on each entry are coming. Nothing is collected here yet.