Model Releases

GPT-5.5 dominates $1,500 LLM hacking test while Gemini refuses to even try

A security researcher spent 1,500 running 13+ AI models against a deliberately vulnerable app, with GPT-5.5 achieving a 70% solve rate while Gemini refused to engage almost entirely. The test app cont

DGX agentreddit
model-releasesr-chatgpt

A security researcher spent $1,500 running 13+ AI models against a deliberately vulnerable app, with GPT-5.5 achieving a 70% solve rate while Gemini refused to engage almost entirely. The test app contained a real-world security vulnerability involving exposed Firebase credentials, and each model was allocated a $10 budget and two hours per run. DeepSeek V4 Pro solved the challenge at just $0.62 per attempt, demonstrating superior cost efficiency despite a lower overall solve rate.

Source: r/ChatGPT | 2026-06-04

Loading related sources…