Additions
1 to 20 of 58, newest first
-
tool
Claude Code
2026-10-09
The way everything on this system gets built, with every session routed through a local proxy so the engine can be a home model or the Claude API. The skills ecosystem (SKILL.md...
-
model
Gemma 4 26B A4B
2026-10-09
The primary model on this system and the right one for a two-card home box. Its mixture-of-experts shape (26B total, about 4B active) gives 65 tokens a second at Q4 with 24K...
-
tool
llama.cpp and llama-server
2026-10-09
The inference server for everything local here, GPU and CPU alike. Native llama-server replaced the Python wrapper in July 2026 and made grammar output nearly free. Watch the slot...
-
practice
A diagram should argue a structure, and the agent should check its own render
2026-10-09
Map visual structure to conceptual structure: fan-out for one-to-many, a timeline for sequence, convergence for aggregation, never a uniform grid of boxes. Render the result and...
-
practice
Grammar-constrained output is nearly free on native llama-server
2026-10-09
On the old Python-wrapped server a grammar cost a flat 3.7× per token. On native llama-server it is gone: 65 tokens a second constrained, the same as free text. Do not avoid...
-
practice
Handle the empty 200, not just the connection error
2026-10-09
A primary that is up but contended returns 200 with an empty completion, never a transport error. A fallback that only fires on connection failure never engages. Retry the empty...
-
practice
Health-check the model, not the process
2026-10-09
A service can be active while its model is still loading, and a metadata API can answer 200 with a dead backend. Probe something that actually runs inference, and expect one to two...
-
practice
Local tiers only: no silent cloud escalation
2026-10-09
If every local tier fails, the step fails and retries next run. It never quietly escalates to a cloud API. Using the Claude API is a deliberate per-session choice, not a fallback.
-
practice
Match response_format to the server that is actually running
2026-10-09
Native llama-server takes {type: json_schema, json_schema: {name, schema, strict}}; the retired llama-cpp-python server took {type: json_object, schema}. Each ignores the other's...
-
practice
max_tokens covers the reasoning too
2026-10-09
With thinking on, a small max_tokens sends the whole allowance to reasoning and returns an empty answer that reads like a parse failure. Budget for both channels or turn thinking...
-
practice
On Ollama, chat_template_kwargs is silently dropped; use the native API with think false
2026-10-09
Ollama's OpenAI-compatible endpoint ignores chat_template_kwargs, so a thinking model keeps thinking. Call the native /api/chat with "think": false instead.
-
practice
A 6 GB card holds one model: do not be the caller that swaps it
2026-10-09
A photo read evicts the text helper and the next text call waits about 12 seconds for the reload. Keep the model tag in config, send no keep_alive and no num_ctx, set the timeout...
-
practice
One inference slot is one caller for the whole system
2026-10-09
The default llama-server runs one slot, so every consumer queues behind every other and the server looks healthy throughout. Check /slots before committing an interactive user to...
-
practice
Parse JSON candidates from the end
2026-10-09
Reasoning models leave rejected drafts in their output. The first JSON object is usually not the answer. Strip think blocks, then walk candidates backwards and take the last one...
-
practice
Pick the host before you stream, and prewarm
2026-10-09
Streaming cannot retry mid-response. Check /health first, then send a one-token non-streamed request to pay the prefill, because the shared server sometimes aborts a streamed...
-
practice
Run a review pass before showing code
2026-10-09
Agents write quickly and review poorly by default. A simplify or code-review skill that runs before the code is presented means what you see is the second draft. Put the project's...
-
practice
Treat schema changes as code: branch, review, merge
2026-10-09
Database decisions made on day one are the hardest to undo on day 365. Make every schema change a branch with a deploy request, select the columns a query needs rather than...
-
practice
A JSON schema constrains the answer, not the thinking
2026-10-09
A response_format guarantees valid JSON only if the model produces content at all. A model that finds the input ambiguous can spend its whole budget reasoning and return an empty...
-
practice
Read other people's sites like a patient person, not a bulk scraper
2026-10-09
Space requests out with jitter, look like an ordinary browser session, arrive the way a reader would, take only what is needed, and never bulk-fetch. A bot-protection page is a...
-
practice
Automated security testing runs only against systems you own, with explicit authorization, never production
2026-10-09
Any agent that executes real attacks against an application must confirm authorization before every run, be scoped to development or staging targets, and keep its tooling in...