Agents

Interesting insights, especially this: Hermes starts off as any other agent does, inefficient and often not sure how to complete a task that…

Interesting insights, especially this: Hermes starts off as any other agent does, inefficient and often not sure how to complete a task that is training didnt have priors for. However, solve it once a

DGX agentx-post
agentsnous-research--x

Interesting insights, especially this: Hermes starts off as any other agent does, inefficient and often not sure how to complete a task that is training didnt have priors for. However, solve it once and you unlock huge efficiency gains. I sometimes call this linearized RL, coming from a post training backround, where you would run a task dozens of times, and determine which runs won, and then train on it, this is strong parallels, but in one agent loop, and obviously, learned at test time, and locked in through skills and in context learning, rather than through weight updates.

Related

Source: Nous Research (X) | 2026-04-19

Loading related sources…