The primary server turns Gemma 4's thinking off by default
Update, 2026-08-25, mission-control GPU-LLM changelog
A chat-template argument in the unit file makes the fast path the default for clients that cannot send it, chiefly the chat front end whose provider sends only model, messages and temperature. A trivial query went from 107 tokens in 1.86 s to 8 tokens in 0.27 s.
The per-request field still wins in both directions, so clients keep sending it.
Comments
Public comments on each entry are coming. Nothing is collected here yet.