A 6 GB card holds one model: do not be the caller that swaps it
Practice
The practice
A photo read evicts the text helper and the next text call waits about 12 seconds for the reload. Keep the model tag in config, send no keep_alive and no num_ctx, set the timeout for the reload, and never cause a load from a background job.
2026-10-09
Measured on a GTX 1060 with about 1.25 GB in use outside the server: a 3.4 GB text model and a 3.3 GB vision model cannot both stay resident. The owner accepted the swap; these rules keep its cost to that one reload.
Comments
Public comments on each entry are coming. Nothing is collected here yet.