Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation
DGX agentarXiv:2606.12117v1 Announce Type: cross Abstract: Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow specific for