Can you explain LLM behavior? If you can predict how the output changes with a different prompt, you're on the right track. Two recent papers used Tinker to test applications of counterfactual simulatability for interpretability:
https://arxiv.org/abs/2602.20710
https://www.alphaxiv.org/pdf/2608.16747
Tinker used to test counterfactual simulatability for LLM interpretability
AISummary
Tinker, the platform from @tinkerapi, supported two recent papers testing counterfactual simulatability as a way to interpret LLM behavior. The core idea is that understanding a model means predicting how its output changes when the prompt changes, with causes ranging from specific words to abstract properties such as a user's angry tone.
Post on XView on X
@tinkerapi
The causes vary widely: specific words in the prompt, quirks of the model (a misremembered fact about an actor), or abstract properties (a user's angry tone).
Source: Tinker · x.comPublished · added here
