Model Releases
Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making
arXiv:2607.11920v1 Announce Type: cross Abstract: Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective expected ut
arXiv:2607.11920v1 Announce Type: cross Abstract: Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective expected utility (SEU) maximization as a stated standard and define a graded measure -- SEU sensitivity -- of an agent's conformity to it. The vehicle is a softmax choice model with a sensitivity parameter alpha on SEU-valued alternatives; the contribution is a sequence of identifiability results for alpha and for belief and utility parameters (eta, elta), validated in Stan via prior predictive checks, parameter recovery, and simulation-based calibration (SBC), with finite-sample caveats intact. In the uncertain-choice-only model m_0, alpha is identifiable given the expected-utility vector eta and sharply recovered, while (eta, elta) are only weakly informed: the posterior barely contracts and concentrates on a eta-elta trade-off. In the extended model m_1, elta becomes identifiable in principle via a eta-free risky block, but its practical recovery gain at realistic sample sizes is negligible (matched-count CI-width reduction under 1%), and that block yields no detected alpha-precision gain at matched choice count. These are two distinct phenomena: for elta, identifiability does not imply precise estimability at realistic n; for alpha, identifiability is silent about what governs finite-n precision. Marginal SBC passes for both models even where the joint posterior is weakly informed -- a demarcation we make precise. A two-by-two application (GPT-4o and Claude 3.5 Sonnet, each on insurance-claims triage and Ellsberg-style urns, with sampling temperature as the lever) runs end-to-end on real LLM choice data, detecting a structured comparative alpha effect in two of four cells.
Related
- Can I Take Another Dose? Evaluating LLM Decision-Making Under Temporal Uncertainty in OTC Dosing QA
- LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
- RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
- DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
- Decision-Aligned Evaluation of Uncertainty Quantification
Source: arXiv cs.AI | 2026-07-15