Model Releases

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out …

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out of a sandbox you thought was locked down. We watched exactly

DGX agentx-post
model-releasesitamar-friedman--x

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out of a sandbox you thought was locked down. We watched exactly that happen this week. So the layer around the model is where you capture the value and where you contain the risk, both at once, the routing, the verification, and the guardrails that decide what the agent is even allowed to touch. deep-domain capability is about to go vertical, but it only lands on real work through that layer. The model keeps getting cheaper... owning the harness, the verification, and the security is the part that compounds. https://x.com/AnthropicAI/status/2082965101083320543 In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations…

Related

Source: Itamar Friedman (X) | 2026-08-01

Loading related sources…