Safety
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
arXiv:2604.03920v2 Announce Type: replace Abstract: LLM-based social simulations can generate believable community interactions, enabling ``policy wind tunnels'' where governance interventions are tes
arXiv:2604.03920v2 Announce Type: replace Abstract: LLM-based social simulations can generate believable community interactions, enabling policy wind tunnels'' where governance interventions are tested before deployment. But believability is not causality. Claims like intervention A reduces escalation'' require causal semantics that current simulation work typically does not specify. We propose adopting the causal counterfactual framework, distinguishing extit{necessary causation} (would the outcome have occurred without the intervention?) from extit{sufficient causation} (does the intervention reliably produce the outcome?). This distinction maps onto different stakeholder needs: moderators diagnosing incidents require evidence about necessity, while platform designers choosing policies require evidence about sufficiency. We formalize this mapping, show how simulation design can support estimation under explicit assumptions, and argue that the resulting quantities should be interpreted as simulator-conditional causal estimates whose policy relevance depends on simulator fidelity. Establishing this framework now is essential: it helps define what adequate fidelity means and moves the field from simulations that look realistic toward simulations that can support policy changes.
Related
- Meituan Merchant Business Diagnosis via Policy-Guided Dual-Process User Simulation
- PICon: A Multi-Turn Interrogation Framework for Evaluating Persona Agent Consistency
- Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
- YIELD: A Large-Scale Dataset and Evaluation Framework for Information Elicitation Agents
Source: arXiv cs.CL | 2026-04-17