Tutorials
The Costs of Pretending That There Are Data-Generating Probability Distributions in the Social World
arXiv:2407.17395v5 Announce Type: replace Abstract: Machine Learning research, including work promoting fair or equitable algorithms, often relies on the concept of a data-generating probability distr
arXiv:2407.17395v5 Announce Type: replace Abstract: Machine Learning research, including work promoting fair or equitable algorithms, often relies on the concept of a data-generating probability distribution. The standard presumption is that since data points are 'sampled from' such a distribution, one can learn from observed data about this distribution and, thus, predict future data points which are also drawn from it. We argue, however, that such true probability distributions do not exist and that the rhetoric around them is harmful in social settings. We show that alternative frameworks focusing directly on relevant populations rather than abstract distributions are available and leave classical learning theory almost unchanged. Furthermore, we argue that the assumption of true probabilities or data-generating distributions can be misleading and obscure both the choices made and the goals pursued in machine learning practice. Based on these considerations, we suggest avoiding the assumption of data-generating probability distributions in the social world.
Related
- Estimating Joint Interventional Distributions from Marginal Interventional Data
- A Review of Diffusion-based Simulation-Based Inference: Foundations and Applications in Non-Ideal Data Scenarios
- Asymptotic Learning Curves for Diffusion Models with Random Features Score and Manifold Data
- KL Divergence Between Gaussians: A Step-by-Step Derivation for the Variational Autoencoder Objective
Source: arXiv cs.LG | 2026-04-23