> vectormatrix.wiki

vLLM

Tool, inference server, by vLLM project

Assessment

The server to reach for when many users share one big GPU. For a home box with two consumer cards and a need to fall back to CPU, llama.cpp's GGUF path fits better, which is why vLLM stays on the watch list rather than in use.

2026-10-09

Strengths

  • High throughput for many concurrent requests on data-centre GPUs
  • Paged attention and continuous batching, the techniques the big hosts use
  • An OpenAI-compatible server with tool calling and structured output

Limitations

  • Wants a recent NVIDIA card with plenty of memory; consumer cards and CPU fallback are not its home
  • GGUF support is partial; its native formats are safetensors and AWQ or GPTQ quantizations

vLLM is an inference server built for throughput: paged attention keeps the key-value cache dense, continuous batching keeps the GPU busy across many requests, and the result is several times the tokens per second of a naive server when dozens of users are talking at once.[1] It speaks the OpenAI API, including tool calls and JSON-schema output.

Where it shines

A shared GPU server: a team, a product, a batch pipeline with thousands of prompts. The gains come from concurrency, so a single user at a time sees little of them.

Why not at home

Two reasons.[2] Its sweet spot is a modern data-centre card with lots of memory; on two 12 GB consumer cards the simplest path is llama.cpp with a GGUF quantization split across them. And a home setup wants a CPU-only fallback for when the GPUs are lent to video work; that is llama.cpp's territory, not vLLM's.

When to revisit

A single larger card, a workload with real concurrency, or a model that ships only in a format llama.cpp handles poorly.

History

  • 2026-10-09: article written and published.
  • 2026-10-09: entry created.

References

  1. vLLM on GitHub ^
  2. vLLM documentation ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.