Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier
DGX agentarXiv:2608.06571v1 Announce Type: cross Abstract: Deployed vision-language systems often gate their answers on confidence, making confidence robustness relevant to oversight. We study confidence reado