Model Releases
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
arXiv:2608.06108v1 Announce Type: new Abstract: Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, p
arXiv:2608.06108v1 Announce Type: new Abstract: Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce extsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a extbf{PrightarrowErightarrowRrightarrowDrightarrowO} trace: investor extit{Profile}, observable market extit{Events}, investment extit{Reasoning}, executable extit{Decision}, and delayed extit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that PrightarrowErightarrowRrightarrowDrightarrowO should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.
Source: arXiv cs.AI | 2026-08-07