Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?
DGX agentarXiv:2602.13626v3 Announce Type: replace Abstract: The expanding integration of Large Language Models (LLMs) into recommender systems poses critical challenges to evaluation reliability. This paper i