Model Releases

Banger paper from the Qwen team. If you evaluate agents on anything longer than a single session, this one is worth your time. (bookmark it)…

Banger paper from the Qwen team. If you evaluate agents on anything longer than a single session, this one is worth your time. (bookmark it) E-Commerce Bench runs an agent through a simulated 365-day

DGX agentx-post
model-releasesdair-ai--x

Banger paper from the Qwen team. If you evaluate agents on anything longer than a single session, this one is worth your time. (bookmark it) E-Commerce Bench runs an agent through a simulated 365-day year operating several online stores at once. 18 frontier models are scored across seven dimensions and no single model dominates. GPT-5.6 Sol earns the most, growing a 100,000 opening stake into 1,431,425, then ranks 16th of 18 on fraud avoidance and trails Fable 5 on operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads at 416,252, 38% above GLM 5.2 (high), and shows the strongest learning over the horizon by progressively bargaining suppliers down across repeated orders. Paper: https://arxiv.org/abs/2608.30730 Chat with Paper: https://academy.dair.ai/papers/e-commerce-bench-evaluating-llm-agents-on-long-horizon-autonomous-business-opera-2608.30730

Related

Source: DAIR.AI (X) | 2026-09-01

Loading related sources…