Model Releases
TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments
arXiv:2608.14270v1 Announce Type: new Abstract: Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of la
arXiv:2608.14270v1 Announce Type: new Abstract: Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions. Existing time series QA benchmarks mostly rely on fixed snapshots, leaving temporal validity and cutoff-aware evidence use unevaluated. We introduce TimeSage-EV, a live benchmark for agentic time series analysis in evolving environments. It tracks 60 real institutional scenarios across 6 domains, comprising 1,485 scenario-period QA pairs from Feb 2023 to May 2026 and spanning monthly, weekly, daily, and irregular release cadences. At each period, large language model (LLM) agents receive time series data and source reports, while the withheld target release provides ground truth. TimeSage-EV evaluates state identification, data summarization, and outlook reasoning. Experiments with frontier LLM agents and TimeSage-1.0, a novel self-evolving agent with a reusable analytical skill library, reveal significant performance gaps across model tiers and recurring failures in temporal validity, exogenous context use, and adaptation. We release TimeSage-EV as a research resource with monthly updates, code, a leaderboard, and failure-mode analyses.
Related
- TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
- TSAQA: Time Series Analysis Question And Answering Benchmark
- Grammar of the Wave: Towards Explainable Multivariate Time Series Event Detection via Neuro-Symbolic VLM Agents
Source: arXiv cs.AI | 2026-08-17