Skip to content
Read the original: Jan Leike· Published 38/100AI score38/100

Jan Leike calls NLAs a new interpretability tool for LLMs

Original titleI'm really excited about this as a new tool in our interpretability tool kit

AISummary

Jan Leike says he is excited about NLAs as a new tool in Anthropic's interpretability toolkit. The quoted post from Sam Marks describes NLAs as an unsupervised method that converts an LLM's internal state into human-readable text, which he says can advance understanding of model thinking and safety auditing.

Read the original x.com

Source: Jan Leike · x.comPublished · added here