> vectormatrix.wiki

max_tokens covers the reasoning too

Practice

The practice

With thinking on, a small max_tokens sends the whole allowance to reasoning and returns an empty answer that reads like a parse failure. Budget for both channels or turn thinking off.

2026-10-09

One project raised its cap to 4096 for this reason. With thinking off, short structured calls stop at end-of-sequence and never approach the cap.

Comments

Public comments on each entry are coming. Nothing is collected here yet.