Safety
I'm really excited about this as a new tool in our interpretability tool kit
I'm really excited about this as a new tool in our interpretability tool kit In a new paper, we present NLAs, an unsupervised method for converting an LLM's internal state into human-readable text. I'
I'm really excited about this as a new tool in our interpretability tool kit In a new paper, we present NLAs, an unsupervised method for converting an LLM's internal state into human-readable text. I've personally been astonished by our results. I think NLAs substantively advance our ability to understand what LLMs are thinking and audit them for safety
Source: Jan Leike (X) | 2026-05-07