Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute
DGX agentarXiv:2608.09351v1 Announce Type: cross Abstract: Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deplo