> vectormatrix.wiki

Gemini API

Provider, cloud llm, by Google

Assessment

Wired as a third engine and rarely chosen: the local model covers most days and Claude the rest. Its long context and multimodal input are the reasons to pick it when they matter; otherwise it is a spare.

2026-10-09

Strengths

  • Very long contexts and native multimodal input: documents, images, audio, video in one request
  • A generous free tier for experiments
  • Fast smaller models for high-volume, low-stakes calls

Limitations

  • Rate limits and quota behaviour change often
  • A third provider to keep wired when two already cover the need

Google's Gemini API serves the Gemini model family, notable for contexts of a million tokens and for taking text, images, audio and video in one request.[1] It has a free tier sized for development and a pay-as-you-go tier above it.

What it is for

Whole-document and whole-video questions, where the context length does the work. Experiments, where the free tier removes the cost question. A second cloud opinion when one is wanted.

Why it is rarely used

A setup with a capable home model for daily work and Claude for the hard cases has little left for a third engine to do, and each extra provider is another key, another quota and another set of quirks to keep working.[2] It stays wired for the day a task fits it.

Things to know

  • Quotas and rate limits have changed several times; read the current table before relying on it in a scheduled job.
  • The same egress rule applies: nothing goes to it without a reason.

History

  • 2026-10-09: article written and published.
  • 2026-10-09: entry created.

References

  1. Google AI for Developers ^
  2. Gemini API documentation ^

Comments

Public comments on each entry are coming. Nothing is collected here yet.