One inference slot is one caller for the whole system
Practice
The practice
The default llama-server runs one slot, so every consumer queues behind every other and the server looks healthy throughout. Check /slots before committing an interactive user to the queue; if you are the batch job, use a smaller tier or accept that you are blocking a person.
2026-10-09
A chat front end that proxies the same server counts as a consumer of the slot too. A project that lists it as a fallback has no fallback: both tiers are the same process.
Comments
Public comments on each entry are coming. Nothing is collected here yet.