From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
arXiv:2608.09925v1 Announce Type: cross Abstract: Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of p