Safety
Semiparametric Double Reinforcement Learning with Applications to Long-Term Causal Inference
arXiv:2501.06926v5 Announce Type: replace-cross Abstract: Double reinforcement learning (DRL) provides efficient off-policy inference for policy values in nonparametric Markov decision processes (MDPs
arXiv:2501.06926v5 Announce Type: replace-cross Abstract: Double reinforcement learning (DRL) provides efficient off-policy inference for policy values in nonparametric Markov decision processes (MDPs), but fully nonparametric estimators can be unstable when intertemporal overlap is weak and occupancy ratios are high-dimensional. This limitation is especially relevant for long-term causal inference from randomized experiments: randomization ensures overlap in treatment assignment, but not over future state trajectories induced by continued intervention use. We develop semiparametric DRL for continuous linear functionals of the infinite-horizon Q-function. Rather than impose linear MDP structure on the reward and transition laws, we place working semiparametric restrictions on the Q-function itself, the solution of the discounted Bellman equation. When correct, these restrictions can improve efficiency relative to unrestricted DRL while allowing rich, possibly infinite-dimensional models. To avoid relying on correct specification, we define the estimand through weighted Bellman-residual minimization. The resulting projection target remains meaningful under misspecification and recovers the original functional under correct specification. For this class of parameters, we derive efficient influence functions and efficiency bounds, construct model-robust automatically debiased estimators, and develop minimax criteria for estimating the Q- and Riesz functions. Under correct specification, optimally weighted versions attain the semiparametric efficiency bound in the restricted model.
Source: arXiv cs.LG | 2026-08-25