Safety
APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent…
APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent about it). Especially on cyber-security, they give a false
APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent about it). Especially on cyber-security, they give a false impression of control and safety whereas in reality they massively increase the risks because they create asymmetry of capabilities and much easier/broader use than open-source model weights even from non-technical people. Anthropic's Mythos has been accessed by a small group of unauthorized users, raising questions about control of the AI model https://www.bloomberg.com/news/articles/2026-04-21/anthropic-s-mythos-model-is-being-accessed-by-unauthorized-users?taid=69e7f03a7728b40001f5a0b0&utm_campa…
Related
- Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs
- Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
- Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion
- Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
Source: Clem Delangue (X) | 2026-04-21