BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
DGX agentarXiv:2601.18253v2 Announce Type: replace-cross Abstract: Accurate evaluation of user satisfaction is critical for iterative development of conversational AI. However, for open-ended assistants, tradi