Research

ClawBench: Can AI Agents Complete Everyday Online Tasks? 153 tasks, 144 live websites, best model at 33.3% [R]

ClawBench is a benchmark of 153 everyday web tasks spanning 144 live platforms across 15 categories — from completing purchases and booking appointments to submitting job applications. Unlike existing

DGX agentreddit
researchr-machinelearning

ClawBench is a benchmark of 153 everyday web tasks spanning 144 live platforms across 15 categories — from completing purchases and booking appointments to submitting job applications. Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and challenges of real-world web interaction. Frontier models such as Claude Sonnet 4.6 and GPT-5.4 score 65–75% on traditional web benchmarks but only 33.3% and 6.5%, respectively, on ClawBench, highlighting a significant gap between current AI agent capabilities and real-world task completion.

Related

Source: r/MachineLearning | 2026-04-14

Loading related sources…