Safety
Adversarial Data Modeling in Epidemiology
arXiv:2602.20134v2 Announce Type: replace-cross Abstract: Epidemiological models increasingly rely on crowdsourced, self-reported behavioral data such as vaccination status, mask usage, and social dis
arXiv:2602.20134v2 Announce Type: replace-cross Abstract: Epidemiological models increasingly rely on crowdsourced, self-reported behavioral data such as vaccination status, mask usage, and social distancing adherence. This data, however, is not passively sampled but instead strategically reported, making it a canonical case of adversarial input to a data mining pipeline. Individuals misreport for various reasons, e.g., to avoid penalties, to access benefits, or to express distrust in public health authorities. We introduce a data-modeling framework that casts the interaction between the population and a public health authority as a signaling game. This approach provides both a generative model of strategically-corrupted behavioral data and a mechanism for the receiver to recover reliable signal from it. Individuals (senders) choose how to report their behaviors, while the public health authority (receiver) updates their epidemiological model(s) based on potentially distorted signals, and modifies its trust in incoming reports accordingly. Focusing on deception around masking and vaccination, we characterize analytically game equilibrium outcomes as distinct regimes of data corruption, and evaluate the degree to which deception can be tolerated while maintaining epidemic control through policy interventions. In large scale simulations, our results show that even under pervasive dishonesty in pooling equilibria, well-designed sender and receiver strategies can still maintain effective epidemic control. Real-world validation further shows that behavioral distortions often exhibit structured patterns rather than arbitrary noise. This work advances the understanding of adversarial data in epidemiology and offers tools for designing more robust public health models in the presence of strategic user behavior.
Related
- Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
- Assessing Trustworthiness of AI Training Dataset using Subjective Logic -- A Use Case on Bias
- When simulations look right but causal effects go wrong: Large language models as behavioral simulators
Source: arXiv cs.AI | 2026-08-19