Grammar-constrained output is nearly free on native llama-server
Practice
The practice
On the old Python-wrapped server a grammar cost a flat 3.7× per token. On native llama-server it is gone: 65 tokens a second constrained, the same as free text. Do not avoid schemas for speed any more.
2026-10-09
The historical measurements live in a performance note in the project that found it. The lesson is that this kind of tax belongs to the server, not the model.
Comments
Public comments on each entry are coming. Nothing is collected here yet.