Model Releases
Nice benchmark to measure agentic e-commerce capabilities. They ran an agent for one simulated year of e-commerce operations and it ends up …
Nice benchmark to measure agentic e-commerce capabilities. They ran an agent for one simulated year of e-commerce operations and it ends up with 27.3% of the money a human makes. MerchantBench is a 36
Nice benchmark to measure agentic e-commerce capabilities. They ran an agent for one simulated year of e-commerce operations and it ends up with 27.3% of the money a human makes. MerchantBench is a 365-day order-level simulation grounded in 98,843 real product records with 26 tools for agent interaction. Agents handle product sourcing, listing and pricing control, cash-flow management, and feedback that arrives at wildly different delays. Scoring runs on cumulative net assets, so incoherence compounds rather than averaging out. Eight LLMs across two agent frameworks, 48 runs of 365 simulated days each, and the best configuration still lands far under the human baseline. Bounded tasks with immediate success criteria have been flattering agents that cannot hold a plan for a month. Paper: https://arxiv.org/abs/2607.28956 Track more trending AI papers in our academy: https://academy.dair.ai/
Source: DAIR.AI (X) | 2026-08-03