> vectormatrix.wiki

LiteLLM

Tool, llm proxy, by BerriAI

Assessment

The translation layer that lets one proxy speak to llama-server, Ollama, Anthropic and Gemini with the same code. Worth it for a switchable-engine setup; read its provider pages, because each backend has a quirk it does not hide.

2026-10-09

Strengths

  • One client shape over a hundred providers, local and cloud
  • Routing, retries and fallbacks between models in configuration
  • Drop-in for code written against the OpenAI SDK

Limitations

  • Provider quirks leak through: a local server's special fields need pass-through
  • Its own defaults can add wasted connection attempts, for example to a localhost Ollama

LiteLLM is a Python library and proxy that presents many model providers behind one interface shaped like the OpenAI API.[1] A request for ollama/qwen2.5-coder:3b, openai/local-big against a llama-server base URL, or anthropic/claude-sonnet-5-5 goes through the same call, and routing rules can send a request to a fallback when the first choice fails.

What it is for

A proxy that lets a coding agent switch engines per session: a home model most days, a cloud model when the task needs it, with the engine chosen in configuration rather than in code.

Things to know

  • Pass-through fields. A local server's special fields, such as the one that turns a model's thinking off, must be forwarded explicitly; the generic layer does not know them.
  • Set the Ollama host. Without OLLAMA_HOST the Ollama provider tries localhost first and burns two refused connections before using the configured base URL.[2]
  • Tool definitions and templates. Some local models cannot parse the tool-call template a provider injects; disabling local tools for that model family is a configuration choice, not a bug.
  • Fallbacks in the proxy are deliberate per-session choices here, never silent escalation from a local model to a cloud one.

History

  • 2026-10-09: article written and published.
  • 2026-10-09: entry created.

References

  1. LiteLLM on GitHub ^
  2. LiteLLM documentation ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.