Model Releases
ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM
ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50
ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50% on these assessments. The benchmark appears designed to measure the ability of large language models to autonomously handle real-world IT operations and administration tasks. This research highlights gaps in AI agent capabilities for enterprise IT environments despite advances in frontier model performance.
Source: Hugging Face | 2026-05-27