When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
DGX agentarXiv:2602.06286v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings where good decisions require forming beliefs over the probability of