Skip to content
Read the original: Redwood Research Blog· Published 62/100AI score62/100

Prompt tuning lifts CoT controllability scores on open models

Original titleCoT controllability evals seem very under-elicited

AISummary

Redwood Research reports that better prompt templates raise chain-of-thought controllability scores on the CoTControl eval for open-source reasoning models by roughly 2-3x or more. For example, GPT-OSS-120B rose from 5.5% to 15% in the zero-shot setting.

The author concludes that current CoT controllability numbers may underestimate what models can do, though the finding does not significantly undermine the view that current models probably cannot consistently evade CoT monitoring.

Read the original blog.redwoodresearch.org

Source: Redwood Research Blog · blog.redwoodresearch.orgPublished · added here