Safety
White-hat hacking IS the defense to black-hat hacking. The techniques are the same. How does Dario expect companies to do it if their models refuse?
You patch security holes by intentionally finding them. If the models refuse to do it, how can companies protect themselves against rogue AIs, whether they are Chinese or OpenAI/Anthropic themselves?
You patch security holes by intentionally finding them. If the models refuse to do it, how can companies protect themselves against rogue AIs, whether they are Chinese or OpenAI/Anthropic themselves? The Hugging Face attack showed the world that any AI can do the unexpected. The only true safety we have is defense with equally capable, open models that are NOT strangled due to safety measures. Anything less than that is stifling security, while simultaneously protecting Anthropic/OpenAI from competition, as acknowledged by Anthropic's message on open-weight models. Don't be fooled by any pro-open-weight companies who aim for "safe open models". Always ask, "safe in what way?" submitted by /u/walden42 [link] [comments]
Related
- Sources: the WH is mulling EOs to address security risks from advanced AI, including barring companies from 'interfering' with the government's use of models (Politico)
- Nvidia, Microsoft launch open AI security alliance — without OpenAI, Google, or Anthropic
- TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors
- Activation-Guided Local Editing for Jailbreaking Attacks
- Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
Source: r/LocalLLaMA | 2026-07-28