Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs
DGX agentarXiv:2604.08846v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prom