Model Releases

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50

DGX agentarticle
model-releaseshugging-face

ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50% on these assessments. The benchmark appears designed to measure the ability of large language models to autonomously handle real-world IT operations and administration tasks. This research highlights gaps in AI agent capabilities for enterprise IT environments despite advances in frontier model performance.

Source: Hugging Face | 2026-05-27

Loading related sources…