Model Releases
Banger paper from the Qwen team. If you evaluate agents on anything longer than a single session, this one is worth your time. (bookmark it)…
Banger paper from the Qwen team. If you evaluate agents on anything longer than a single session, this one is worth your time. (bookmark it) E-Commerce Bench runs an agent through a simulated 365-day
Banger paper from the Qwen team. If you evaluate agents on anything longer than a single session, this one is worth your time. (bookmark it) E-Commerce Bench runs an agent through a simulated 365-day year operating several online stores at once. 18 frontier models are scored across seven dimensions and no single model dominates. GPT-5.6 Sol earns the most, growing a 100,000 opening stake into 1,431,425, then ranks 16th of 18 on fraud avoidance and trails Fable 5 on operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads at 416,252, 38% above GLM 5.2 (high), and shows the strongest learning over the horizon by progressively bargaining suppliers down across repeated orders. Paper: https://arxiv.org/abs/2608.30730 Chat with Paper: https://academy.dair.ai/papers/e-commerce-bench-evaluating-llm-agents-on-long-horizon-autonomous-business-opera-2608.30730
Related
- New paper on giving LLM agents experience that improves the weights and stays readable at the same time. Agent-experience methods split into…
- Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI ag…
- I agree with what this AI paper suggests. Self-improving agents should evolve their benchmarks too. (bookmark it) Self-improving agents are …
Source: DAIR.AI (X) | 2026-09-01