Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
DGX agentarXiv:2606.00093v1 Announce Type: new Abstract: Validating an LLM judge against human annotations usually means reporting several agreement statistics: accuracy, precision, recall, F_1, Cohen's kappa,