Mechanisms of Introspective Awareness
DGX agentarXiv:2603.21396v2 Announce Type: replace Abstract: Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept
Knowledge catalogue
arXiv:2603.21396v2 Announce Type: replace Abstract: Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept
arXiv:2604.09326v1 Announce Type: cross Abstract: Ensuring safety and reliability in human-robot interaction (HRI) requires the timely detection of unexpected events that could lead to system failures