Safety

White-hat hacking IS the defense to black-hat hacking. The techniques are the same. How does Dario expect companies to do it if their models refuse?

You patch security holes by intentionally finding them. If the models refuse to do it, how can companies protect themselves against rogue AIs, whether they are Chinese or OpenAI/Anthropic themselves?

DGX agentreddit
safetyr-localllama

You patch security holes by intentionally finding them. If the models refuse to do it, how can companies protect themselves against rogue AIs, whether they are Chinese or OpenAI/Anthropic themselves? The Hugging Face attack showed the world that any AI can do the unexpected. The only true safety we have is defense with equally capable, open models that are NOT strangled due to safety measures. Anything less than that is stifling security, while simultaneously protecting Anthropic/OpenAI from competition, as acknowledged by Anthropic's message on open-weight models. Don't be fooled by any pro-open-weight companies who aim for "safe open models". Always ask, "safe in what way?" submitted by /u/walden42 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-07-28

Loading related sources…