Safety

I'm really excited about this as a new tool in our interpretability tool kit

I'm really excited about this as a new tool in our interpretability tool kit In a new paper, we present NLAs, an unsupervised method for converting an LLM's internal state into human-readable text. I'

DGX agentx-post
safetyjan-leike--x

I'm really excited about this as a new tool in our interpretability tool kit In a new paper, we present NLAs, an unsupervised method for converting an LLM's internal state into human-readable text. I've personally been astonished by our results. I think NLAs substantively advance our ability to understand what LLMs are thinking and audit them for safety

Source: Jan Leike (X) | 2026-05-07

Loading related sources…