Skip to content
View original post on X: Tinker· 25/100AI score25/100

Tinker highlights training objectives for legible chain-of-thought and interpretability evals

AISummary

Tinker says Hase & Potts convert a model's chain-of-thought into a training objective so a monitor can read it more easily. Karvonen et al. use tested counterfactual outputs to build an interpretability eval. The post notes that counterfactuals do not explain the underlying mechanism, but their predictability is a useful foundation.

Post on XView on X
@tinkerapi

A reply · the post it answers

Hase & Potts turn this into a training objective, making a model's CoT legible to a monitor. Karvonen et al. use the tested output of counterfactuals to build an interpretability eval.

Counterfactuals don't explain the mechanism, but predictability is a good base to build on.

Source: Tinker · x.comPublished · added here