Research
ConvApparel: Measuring and bridging the realism gap in user simulators
ConvApparel is a human-AI conversation dataset and comprehensive evaluation framework designed to quantify the 'realism gap' in LLM-based user simulators and improve the training of robust conversa...
ConvApparel is a human-AI conversation dataset and comprehensive evaluation framework designed to quantify the "realism gap" in LLM-based user simulators and improve the training of robust conversational agents. Its dual-agent data collection protocol — using both "good" and "bad" recommenders — captures a wide spectrum of user experiences enriched with satisfaction annotations, while its validation framework combines statistical alignment, a human-likeness score, and counterfactual validation to test for generalization. Experiments reveal a significant realism gap across all tested simulators, though data-driven simulators outperform prompted baselines — particularly in counterfactual settings where they adapt more realistically to unseen behaviors.
Related
- STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems
- A Goal-Oriented Chatbot for Engaging the Elderly Through Family Photo Conversations
- Free tool I built to score dataset quality (LQS) — feedback welcome [D]
- SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model Training
Source: research