Research
Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
arXiv:2608.24842v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Y
arXiv:2608.24842v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.
Related
- To LLM, or Not to LLM: How Designers and Developers Navigate LLMs as Tools or Teammates
- Generative AI-Based Virtual Assistant using Retrieval-Augmented Generation: An evaluation study for bachelor projects
- Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning
- Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
Source: arXiv cs.AI | 2026-08-26