> vectormatrix.wiki

What's new

Latest additions

  • tool
    Claude Code

    2026-10-09

    The way everything on this system gets built, with every session routed through a local proxy so the engine can be a home model or the Claude API. The skills ecosystem (SKILL.md...

  • model
    Gemma 4 26B A4B

    2026-10-09

    The primary model on this system and the right one for a two-card home box. Its mixture-of-experts shape (26B total, about 4B active) gives 65 tokens a second at Q4 with 24K...

  • tool
    llama.cpp and llama-server

    2026-10-09

    The inference server for everything local here, GPU and CPU alike. Native llama-server replaced the Python wrapper in July 2026 and made grammar output nearly free. Watch the slot...

  • practice
    A diagram should argue a structure, and the agent should check its own render

    2026-10-09

    Map visual structure to conceptual structure: fan-out for one-to-many, a timeline for sequence, convergence for aggregation, never a uniform grid of boxes. Render the result and...

  • practice
    Grammar-constrained output is nearly free on native llama-server

    2026-10-09

    On the old Python-wrapped server a grammar cost a flat 3.7× per token. On native llama-server it is gone: 65 tokens a second constrained, the same as free text. Do not avoid...

  • practice
    Handle the empty 200, not just the connection error

    2026-10-09

    A primary that is up but contended returns 200 with an empty completion, never a transport error. A fallback that only fires on connection failure never engages. Retry the empty...

  • practice
    Health-check the model, not the process

    2026-10-09

    A service can be active while its model is still loading, and a metadata API can answer 200 with a dead backend. Probe something that actually runs inference, and expect one to two...

  • practice
    Local tiers only: no silent cloud escalation

    2026-10-09

    If every local tier fails, the step fails and retries next run. It never quietly escalates to a cloud API. Using the Claude API is a deliberate per-session choice, not a fallback.

  • practice
    Match response_format to the server that is actually running

    2026-10-09

    Native llama-server takes {type: json_schema, json_schema: {name, schema, strict}}; the retired llama-cpp-python server took {type: json_object, schema}. Each ignores the other's...

  • practice
    max_tokens covers the reasoning too

    2026-10-09

    With thinking on, a small max_tokens sends the whole allowance to reasoning and returns an empty answer that reads like a parse failure. Budget for both channels or turn thinking...

More additions

Latest from the feed

More from the feed