Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models
DGX agentarXiv:2608.01624v1 Announce Type: new Abstract: Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable c