Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
DGX agentarXiv:2607.02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and