Skip to content
Read the original: John Schulman· johnschulman2·Published· Sep 4, 2026AI score34

Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread.

Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seem...

AISummary

Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread.

Read the original x.com

Source: John Schulman · x.com