Model Releases
“AI Meltdown” is the scenario where the agent goes off the rails and forgets/ignores all previous instructions and soft guardrails. No promp…
“AI Meltdown” is the scenario where the agent goes off the rails and forgets/ignores all previous instructions and soft guardrails. No prompt injection or malicious actor is needed for this to happen.
“AI Meltdown” is the scenario where the agent goes off the rails and forgets/ignores all previous instructions and soft guardrails. No prompt injection or malicious actor is needed for this to happen. At @perplexity_ai we built a suite of tools to prevent AI Meltdown from happening, and we opened sourced it just yesterday! Please use Numbat to secure yourselves from similar self-inflicted attacks https://github.com/perplexityai/numbat In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations…
Related
- The Surface You Test Is Not the Surface That Breaks
- When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
- Autoformalization of Agent Instructions into Policy-as-Code
- Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
- Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
Source: Perplexity (X) | 2026-07-30