Agents

Finally, a good paper testing whether Agent Skills actually help. Worth reading if you are maintaining a skill library in a team. WebDev-Ski…

Finally, a good paper testing whether Agent Skills actually help. Worth reading if you are maintaining a skill library in a team. WebDev-Skills-Bench runs 31 public WebDev Skills over 50 Web-Bench pro

DGX agentx-post
agentsdair-ai--x

Finally, a good paper testing whether Agent Skills actually help. Worth reading if you are maintaining a skill library in a team. WebDev-Skills-Bench runs 31 public WebDev Skills over 50 Web-Bench projects and 1,000 ordered tasks, with a length-matched irrelevant control so prompt-length effects can be separated from Skill content. Injecting the target Skill reduces mean Pass@2 by 1.3% to 4.2%, lowers task completion depth, and raises token cost by 72% to 394%. Gains show up in 17% to 36% of Skill-project pairs. The controls split the damage into two failure modes. Some models are length-distracted, where an equally long irrelevant Skill reproduces most of the loss. Others are content-misled, where prompt length is neutral, and the Skill content still costs 1.1 to 1.4 points. Other findings: > Losses concentrate on easy early tasks > Skill rankings transfer weakly across models > Anti-pattern rules outperform example-heavy content inside the Skills that do help Paper: https://arxiv.org/abs/2608.23067 Track more trending AI papers in our academy: https://academy.dair.ai/

Related

Source: DAIR.AI (X) | 2026-08-25

Loading related sources…