Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
DGX agentarXiv:2604.08970v1 Announce Type: cross Abstract: We study predictive multilingual evaluation: estimating how well a model will perform on a task in a target language when direct benchmark results are