Model Releases

GROK 4.5 LEADS ON REAL PROFESSIONAL WORK BENCHMARK New data from Snorkel shows Grok 4.5 outperforming other frontier models on real-world pr…

GROK 4.5 LEADS ON REAL PROFESSIONAL WORK BENCHMARK New data from Snorkel shows Grok 4.5 outperforming other frontier models on real-world professional tasks. On their GDPval+ benchmark (expert-created

DGX agentx-post
model-releaseselon-musk--x

GROK 4.5 LEADS ON REAL PROFESSIONAL WORK BENCHMARK New data from Snorkel shows Grok 4.5 outperforming other frontier models on real-world professional tasks. On their GDPval+ benchmark (expert-created workplace reasoning tasks across the economy): • Grok 4.5: 29% mean pass rate • GPT 5.5: 22% • Claude Opus 4.8: 21% Grok 4.5 showed particularly strong gains in demanding areas like legal work, education, healthcare, and QA analysis. This lines up with xAI’s focus on building models that excel at practical, agentic work rather than just synthetic benchmarks. While general intelligence leaderboards still see tight competition at the very top, Grok 4.5 is delivering some of the strongest results on actual professional deliverables right now. Grok Build improves almost every day

Source: Elon Musk (X) | 2026-07-10

Loading related sources…