Gemma 4 on llama-server reads images: the vision adapter is resident
Update, 2026-09-24, mission-control GPU-LLM changelog
The OCR service now sends photos to the primary model itself: an OpenAI image_url content part, four or five short requests per photo. The adapter costs about 1.5 GB on the slow card and leaves decode speed unchanged. The first image aborted the server until the micro-batch size was raised above the image token count.
Keep the server's micro-batch at or above the image token limit; at 256 against 280 the first photo took the whole server down.
Comments
Public comments on each entry are coming. Nothing is collected here yet.