LiteLLM
Tool, llm proxy, by BerriAI
Assessment
The translation layer that lets one proxy speak to llama-server, Ollama, Anthropic and Gemini with the same code. Worth it for a switchable-engine setup; read its provider pages, because each backend has a quirk it does not hide.
2026-10-09
Strengths
- One client shape over a hundred providers, local and cloud
- Routing, retries and fallbacks between models in configuration
- Drop-in for code written against the OpenAI SDK
Limitations
- Provider quirks leak through: a local server's special fields need pass-through
- Its own defaults can add wasted connection attempts, for example to a localhost Ollama
LiteLLM is a Python library and proxy that presents many model providers behind one interface shaped like the OpenAI API.[1] A request for ollama/qwen2.5-coder:3b, openai/local-big against a llama-server base URL, or anthropic/claude-sonnet-5-5 goes through the same call, and routing rules can send a request to a fallback when the first choice fails.
What it is for
A proxy that lets a coding agent switch engines per session: a home model most days, a cloud model when the task needs it, with the engine chosen in configuration rather than in code.
Things to know
- Pass-through fields. A local server's special fields, such as the one that turns a model's thinking off, must be forwarded explicitly; the generic layer does not know them.
- Set the Ollama host. Without
OLLAMA_HOSTthe Ollama provider tries localhost first and burns two refused connections before using the configured base URL.[2] - Tool definitions and templates. Some local models cannot parse the tool-call template a provider injects; disabling local tools for that model family is a configuration choice, not a bug.
- Fallbacks in the proxy are deliberate per-session choices here, never silent escalation from a local model to a cloud one.
History
- 2026-10-09: article written and published.
- 2026-10-09: entry created.
Comments
Public comments on each entry are coming. Nothing is collected here yet.