Model Releases

// Memory is what breaks long-horizon agents // It's undeniable how important memory/recall is for long-horizon tasks. If you are building f…

// Memory is what breaks long-horizon agents // It's undeniable how important memory/recall is for long-horizon tasks. If you are building for long-horizon tasks, this is a great read. (bookmark it) T

DGX agentx-post
model-releasesdair-ai--x

// Memory is what breaks long-horizon agents // It's undeniable how important memory/recall is for long-horizon tasks. If you are building for long-horizon tasks, this is a great read. (bookmark it) They set up an LLM agent to run a football club for 20 in-game years, through 26 tools and roughly 340 to 400 decision stops, scored by a deterministic engine with no LLM judge anywhere in the loop. Results: All 15 frontier models survive every horizon while the scripted baselines mostly die out. Neither scale, price, vendor, nor token spend predicts the ranking, and the order only settles late in the run. What separates the top models is managerial behavior, cutting slow-payoff investment near the end and opening contract renewals well before the deadline. Two universal failures were found. No model learns the market's hidden prices from hundreds of rejected bids, and self-managed memory collapses into either an archive that only grows or a plan rewritten every season. Paper: https://arxiv.org/abs/2608.18423 Chat with Paper: https://academy.dair.ai/papers/fm-bench-a-benchmark-for-long-horizon-management-with-competing-agents-2608.18423

Related

Source: DAIR.AI (X) | 2026-08-29

Loading related sources…