Model Releases
ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents
arXiv:2602.10863v2 Announce Type: replace-cross Abstract: Long-horizon reinforcement learning for information seeking agents remains difficult because terminal rewards reveal whether the final answer
arXiv:2602.10863v2 Announce Type: replace-cross Abstract: Long-horizon reinforcement learning for information seeking agents remains difficult because terminal rewards reveal whether the final answer is correct, but not which acquired information enabled it. This difficulty is amplified by text-derived webpage observations, where parsing, truncation, and summarization often produce incomplete and unstable content representations across trajectories. We propose an evidence-centric framework for web agent learning that represents information acquired through tools as identifiable units for comparison across trajectories. In particular, fetched webpages are represented as rendered snapshots, preserving layout and multimodal content as stable content-level observations. Building on these units, we introduce Information-Aware Credit Assignment, a post hoc reward propagation method that estimates turn-level utility scores from rollout success rates and assigns dense rewards to intermediate steps that introduced high-utility information. Integrated with GSPO, our method consistently improves performance on BrowseComp, GAIA, Xbench-DS, and Seal-0. Code and datasets will be released at https://github.com/pc-inno/ICA_MM_deepsearch.
Related
- MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents
- HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
- HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
Source: arXiv cs.AI | 2026-08-26