Agents
AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval
arXiv:2608.24076v1 Announce Type: new Abstract: Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and
arXiv:2608.24076v1 Announce Type: new Abstract: Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and adversarial brittleness. We present AgentWorld, a simulation framework combining (i)Big Five (OCEAN) personality-driven user populations with stateful tool-use environments; (ii)the pass^k consistency metric with structured fault classification, partial-credit scoring, and dual-control handoff verification; (iii)score-thresholded training-data export in six fine-tuning formats; and (iv)an adversarial Risk Analyser that snapshots required-intermediate-state spines, branches Monte-Carlo rollouts under four task-aware perturbation types, and quantifies risk via Delta P / Delta T scoring, Dempster--Shafer evidence fusion, and Shapley attack-category attribution. Three experiments demonstrate the framework: a conversational analytics agent across 10 OCEAN personas (240 evaluator judgments); a customer-support agent across 5 tasks imes 4 persona variants; and adversarial stress-testing of 5 tasks revealing pre-existing trajectory brittleness (V_{min}=0.375 without perturbation) and tool/infrastructure-layer attack dominance (Shapley: 46% system, 38% action). Personality variation surfaces failure modes uniform testing cannot expose---cross-domain leakage, contextual drift, a 0.27-point quality gap, and 50% vs. 100% pass-rate across personas on the same task---while the Risk Analyser quantifies trajectory-level brittleness that pass^k alone cannot measure.
Related
- Efficient Agent Evaluation via Diversity-Guided User Simulation
- Coverage Aware Active Evaluation for Failure Discovery with Paired Systems
- Can LLMs Infer Conversational Agent Users' Personality Traits from Chat History?
Source: arXiv cs.AI | 2026-08-26