Unstable Rankings in Bayesian Deep Learning Evaluation
DGX agentarXiv:2604.23102v1 Announce Type: new Abstract: Standard evaluations of Bayesian deep learning methods assume that metric estimates are reliable, but we show this assumption fails under data scarcity.