Research

All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting

arXiv:2602.17234v2 Announce Type: replace Abstract: Backtesting LLMs on resolved events assumes models reason only from pre-cutoff knowledge, yet pretrained models inevitably leak post-cutoff knowledg

DGX agentpaper
researcharxiv-cs-ai

arXiv:2602.17234v2 Announce Type: replace Abstract: Backtesting LLMs on resolved events assumes models reason only from pre-cutoff knowledge, yet pretrained models inevitably leak post-cutoff knowledge. We introduce a claim-level evaluation framework that decomposes prediction rationales into atomic claims and applies Shapley values to quantify each claim's decision impact, yielding extbf{Shapley-DCLR} (extbf{Shapley}-weighted extbf{D}ecision-extbf{C}ritical extbf{L}eakage extbf{R}ate) -- an interpretable metric measuring what fraction of decision-driving reasoning is contaminated. We further propose extbf{TimeSPEC} (extbf{Time}-extbf{S}upervised extbf{P}rediction with extbf{E}xtracted extbf{C}laims), an inference-time architecture that interleaves temporally-filtered retrieval with claim-level supervision, producing predictions grounded entirely in pre-cutoff evidence. Across three LLMs, the ablation experiments confirm retrieval and supervision are jointly necessary; and a three-task probe further illstrates that the performance cost of temporal enforcement scales with each task's reliance on post-cutoff information.

Source: arXiv cs.AI | 2026-05-26

Loading related sources…