Model Releases
Detecting critical treatment effect bias in small subgroups
arXiv:2404.18905v3 Announce Type: replace-cross Abstract: Randomized trials are considered the gold standard for making informed decisions in medicine, yet they often lack generalizability to the pati
arXiv:2404.18905v3 Announce Type: replace-cross Abstract: Randomized trials are considered the gold standard for making informed decisions in medicine, yet they often lack generalizability to the patient populations in clinical practice. Observational studies, on the other hand, cover a broader patient population but are prone to various biases. Thus, before using an observational study for decision-making, it is crucial to benchmark its treatment effect estimates against those derived from a randomized trial. We propose a novel strategy to benchmark observational studies beyond the average treatment effect. First, we design a statistical test for the null hypothesis that the treatment effects estimated from the two studies, conditioned on a set of relevant features, differ up to some tolerance. We then estimate an asymptotically valid lower bound on the maximum bias strength for any subgroup in the observational study. Finally, we validate our benchmarking strategy in a real-world setting and show that it leads to conclusions that align with established medical knowledge.
Related
- Bias Detection in Emergency Psychiatry: Linking Negative Language to Diagnostic Disparities
- Flow-based Generative Modeling of Potential Outcomes and Counterfactuals
- Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models
- Accurate and Reliable Uncertainty Estimates for Deterministic Predictions Extensions to Under and Overpredictions
Source: arXiv cs.LG | 2026-04-14