Jan Leike calls NLAs a new interpretability tool for LLMs
Original titleI'm really excited about this as a new tool in our interpretability tool kit
AISummary
Jan Leike says he is excited about NLAs as a new tool in Anthropic's interpretability toolkit. The quoted post from Sam Marks describes NLAs as an unsupervised method that converts an LLM's internal state into human-readable text, which he says can advance understanding of model thinking and safety auditing.
Source: Jan Leike · x.comPublished · added here