Additions
21 to 40 of 58, newest first
-
practice
Catch the transport error family and set a short connect timeout
2026-10-09
A stopped service does not always refuse; it can hang until a connect timeout, which is a sibling of the connection error, not a subclass. Catch the family, use about five seconds...
-
practice
Prefer skills that change the default over skills you must remember to invoke
2026-10-09
The test for any agent skill: does it change what the agent produces without constant prompting, or does it just add a command to a menu? Frontend-design changes what a...
-
practice
When an agent uses live or paid data, surface the sources
2026-10-09
Data pulled from search or a specialised index is only as trustworthy as its citation trail. Be specific about which sources a query may use, prefer a cited-answer call when a...
-
practice
Turn the reasoning channel off on Gemma 4, in every request
2026-10-09
Send chat_template_kwargs {enable_thinking: false} with every call, even though the server defaults it off. A grammar-constrained call went from 600 tokens and 8.7 s to 72 tokens...
-
update
The primary server turns Gemma 4's thinking off by default
2026-10-09
A chat-template argument in the unit file makes the fast path the default for clients that cannot send it, chiefly the chat front end whose provider sends only model, messages and...
-
update
Gemma 4 on llama-server reads images: the vision adapter is resident
2026-10-09
The OCR service now sends photos to the primary model itself: an OpenAI image_url content part, four or five short requests per photo. The adapter costs about 1.5 GB on the slow...
-
update
VectorMatrix.wiki seeded from the home system's own records
2026-10-09
First entries: local models and inference servers, the LLM client rules as practices, and the ten agent skills from a March article.
-
provider
Claude (Anthropic API)
2026-10-09
The cloud engine for the work a home model cannot do: large refactors, hard reasoning, long documents. Treat each use as a deliberate choice, never a silent fallback, because the...
-
tool
AnythingLLM
2026-10-09
A good direct chat and retrieval front end for a local server. Remember what it is: every query it makes is a call to the same model behind it, queued with everyone else, and it...
-
tool
ComfyUI
2026-10-09
The engine for everything generative that is not text: image, video, lip sync, 3D. Run it as a service with saved workflows and an API, and accept that it takes the cards when it...
-
tool
Eleventy
2026-10-09
The generator this site is built with. Its data cascade fits a site of structured entries better than the alternatives: a folder is a collection, a file is a record, and a filter...
-
tool
faster-whisper
2026-10-09
The transcriber for quality: large-v3-turbo on a GPU for batch work, and the same library on a CPU as the second home for live calls. Wrap it in whisper.cpp's request shape so...
-
provider
Gemini API
2026-10-09
Wired as a third engine and rarely chosen: the local model covers most days and Claude the rest. Its long context and multimodal input are the reasons to pick it when they matter;...
-
model
Gemma 4 E2B
2026-10-09
Tried as the small helper and set aside. It reasons by default, which the OpenAI-style endpoint cannot switch off, its file does not fit a 6 GB card, and its vision misread a plain...
-
tool
Hugo
2026-10-09
The generator to pick when build speed or a dependency-free install matters most. For a site that is really a set of structured records, Eleventy's data model won, which is the...
-
model
Kokoro TTS
2026-10-09
The best-sounding voice for its size by a distance, and the one voice engine that needs no GPU mode switch because it fits beside a loaded LLM. The right pick for narration and...
-
tool
LiteLLM
2026-10-09
The translation layer that lets one proxy speak to llama-server, Ollama, Anthropic and Gemini with the same code. Worth it for a switchable-engine setup; read its provider pages,...
-
model
LTX-2.3 22B
2026-10-09
The heaviest thing on a two-card box and the only one that makes a face speak convincingly. Worth it for finished pieces, not for experiments. Budget half an hour a clip and give...
-
tool
Ollama
2026-10-09
The easiest way to serve a model on a desktop, especially on Windows. Use it for the helper tier and talk to its native API when a model has a thinking mode. Expect the...
-
tool
Pagefind
2026-10-09
The search for a static site: run it over the output, ship the folder it writes, and nothing is sent anywhere when someone searches. Its lazy-loaded index makes it practical even...