Beyond Memorization: Assessing Semantic Generalization in Large Language Models Using Phrasal Constructions
arXiv:2501.04661v3 Announce Type: replace-cross Abstract: The web-scale of pretraining data has created an important evaluation challenge: to disentangle linguistic competence on cases well-represente