Personalized Benchmarking: Evaluating LLMs by Individual Preferences
DGX agentarXiv:2604.18943v1 Announce Type: new Abstract: With the rise in capabilities of large language models (LLMs) and their deployment in real-world tasks, evaluating LLM alignment with human preferences