When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
DGX agentarXiv:2509.26600v2 Announce Type: replace-cross Abstract: As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inp