vLLM
Tool, inference server, by vLLM project
Assessment
The server to reach for when many users share one big GPU. For a home box with two consumer cards and a need to fall back to CPU, llama.cpp's GGUF path fits better, which is why vLLM stays on the watch list rather than in use.
2026-10-09
Strengths
- High throughput for many concurrent requests on data-centre GPUs
- Paged attention and continuous batching, the techniques the big hosts use
- An OpenAI-compatible server with tool calling and structured output
Limitations
- Wants a recent NVIDIA card with plenty of memory; consumer cards and CPU fallback are not its home
- GGUF support is partial; its native formats are safetensors and AWQ or GPTQ quantizations
vLLM is an inference server built for throughput: paged attention keeps the key-value cache dense, continuous batching keeps the GPU busy across many requests, and the result is several times the tokens per second of a naive server when dozens of users are talking at once.[1] It speaks the OpenAI API, including tool calls and JSON-schema output.
Where it shines
A shared GPU server: a team, a product, a batch pipeline with thousands of prompts. The gains come from concurrency, so a single user at a time sees little of them.
Why not at home
Two reasons.[2] Its sweet spot is a modern data-centre card with lots of memory; on two 12 GB consumer cards the simplest path is llama.cpp with a GGUF quantization split across them. And a home setup wants a CPU-only fallback for when the GPUs are lent to video work; that is llama.cpp's territory, not vLLM's.
When to revisit
A single larger card, a workload with real concurrency, or a model that ships only in a format llama.cpp handles poorly.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.