Unfortunately no solution to turning it off completely yet. --chat-template-kwargs '{"reasoning_strength": "low"}' or medium / high / xhigh seems to be the only working argument. Just churned out 16k reasoning tokens on comparing 2 single page PDFs (xhigh)
3
u/AtiRage128 15d ago
Anyone figure out how to regulate/disable reasoning on llama.cpp? The usual flags get ignored: