Research

Stable Global Weighting of Flow Mixtures using Simplex Exponential Moving Average

arXiv:2607.03809v1 Announce Type: new Abstract: Normalising flows provide a powerful variational family for approximate inference, yet individual architectures often fail to generalise across heteroge

DGX agentpaper
researcharxiv-cs-lg

arXiv:2607.03809v1 Announce Type: new Abstract: Normalising flows provide a powerful variational family for approximate inference, yet individual architectures often fail to generalise across heterogeneous posterior geometries. We revisit mixture-based flow formulations and introduce AMFmbox{-VImbox{-}sEMA}, a two-stage framework featuring a stable global weighting mechanism based on a Simplex Exponential Moving Average (sEMA) update. In Stage1, a heterogeneous set of experts (extsc{RealNVP}, extsc{MAF}, extsc{RBIG}) are trained independently to specialise in distinct structural regimes. In Stage2, expert parameters are frozen and global mixture weights are learned through a temperature-controlled softmax of average log-likelihoods, followed by a smooth EMA update on the probability simplex. This design produces a tractable, data-agnostic gating mechanism (without per-sample gating or gradient backpropagation through weights) that adaptively reallocates capacity while avoiding component collapse. We evaluate the framework on ten posterior benchmarks: six canonical 2D synthetic families (Banana, X-Shaped, Bimodal, Multimodal, Two-moons, Rings) and four real/low-dimensional Bayesian targets (BLR, BPR, Weibull, Real-GMM2), with stronger baselines (extsc{NICE}, extsc{ResFlow}, and EM-Mixing). Comprehensive evaluation covers NLL, KL divergence, Wasserstein-2 distance, and MMD, together with diagnostics of mixture dynamics, hyperparameter sensitivity, and cross-seed robustness. Empirically, AMFmbox{-VImbox{-}sEMA} achieves consistent NLL improvements over its predecessor AMFmbox{-VI} and avoids the catastrophic transport failures of single-flow baselines, while maintaining stable weight trajectories (N_{eff}{>}1.4 on all datasets) with minimal computational overhead.

Source: arXiv cs.LG | 2026-07-07

Loading related sources…