Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
DGX agentarXiv:2606.27226v1 Announce Type: new Abstract: Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments