Safety
QuantFPFlow: Quantum Amplitude Estimation for Fokker--Planck Policy Optimisation in Continuous Reinforcement Learning
arXiv:2605.16429v1 Announce Type: cross Abstract: We introduce extbf{QuantFPFlow}, a reinforcement learning framework that integrates quantum amplitude estimation into the Fokker--Planck~(FP) formulat
arXiv:2605.16429v1 Announce Type: cross Abstract: We introduce extbf{QuantFPFlow}, a reinforcement learning framework that integrates quantum amplitude estimation into the Fokker--Planck~(FP) formulation of stochastic policy optimisation. Classical continuous-space RL agents must estimate the FP partition function Z = int e^{-V(mathbf{x})/D},dmathbf{x} at cost alO(1/arepsilon^{2}); QuantFPFlow replaces this with a Grover-amplified amplitude estimator achieving alO(1/arepsilon) -- a provable quadratic speedup. While the full quantum acceleration requires fault-tolerant hardware, the quantum-inspired classical simulation demonstrated here already exhibits the alO(1/arepsilon) algorithmic structure. The estimated stationary distribution rhostar drives a theoretically grounded exploration bonus Raug = Renv + alphalog(1/rhostar(s)). This bonus steers the agent toward globally optimal regions of multimodal reward landscapes while simultaneously constraining policy variance through FP diffusion matching. On a continuous-control task specifically designed to expose local-optima failure, QuantFPFlow achieves mean reward 1{,}295.7 pm 423.2 versus 1{,}284.0 pm 474.0 for Soft Actor-Critic~(SAC), while discovering the global optimum extbf{10.4,% more frequently} (33.9,% vs. 30.7,%). Policy entropy remains near H(pi)approx 6.5,nats throughout training, whereas SAC collapses to 1.5,nats, confirming that FP diffusion matching actively prevents premature convergence. Dimensionality experiments further show computational scaling of alO(d^{0.35}) for QuantFPFlow versus alO(d^{0.76}) for classical FP estimation.
Source: arXiv cs.AI | 2026-05-19