max_tokens covers the reasoning too
Practice
The practice
With thinking on, a small max_tokens sends the whole allowance to reasoning and returns an empty answer that reads like a parse failure. Budget for both channels or turn thinking off.
2026-10-09
One project raised its cap to 4096 for this reason. With thinking off, short structured calls stop at end-of-sequence and never approach the cap.
Comments
Public comments on each entry are coming. Nothing is collected here yet.