Tsinghua, Oxford and Stanford paper finds LLMs keep reasoning when told not to
Overview
A Tsinghua, Oxford and Stanford paper finds that LLMs still write out reasoning with thinking disabled, especially on open-ended questions.
With thinking off, DeepSeek-V4-Flash wrote out reasoning in 99.9% of open-ended answers.
Forcing answer-only replies on open-ended tasks raised compliance to about 40% across five models, according to the paper, but cut accuracy by about 15 points.
Written by AI from the articles below · updated Oct 10, 12:40 AM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
Rohan Paul@rohanpaul_aiXLLMs keep reasoning out loud even with thinking disabledAIA Tsinghua, Oxford, and Stanford paper finds that LLMs still write out reasoning with thinking turned off, especially on open-ended questions. With thinking disabled, DeepSeek-V4-Flash wrote reasoning in 99.9% of open-ended answers. Forcing answer-only replies on open-ended tasks raised compliance to about 40% across 5 models but cut accuracy by about 15 points.

Heat trend
Not enough continuous observations to show a trend yet.