Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
DGX agentarXiv:2605.00119v1 Announce Type: new Abstract: There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. M