Agents

4/5 Best-of-N: Run one agent config N times in parallel → select the best trajectory. Leverages LLMs’ non-determinism - but hinges on a good…

4/5 Best-of-N: Run one agent config N times in parallel → select the best trajectory. Leverages LLMs’ non-determinism - but hinges on a good eval mechanism (we use an LLM-as-a-Judge). Can increase acc

DGX agentx-post
agentsai21-labs--x

4/5 Best-of-N: Run one agent config N times in parallel → select the best trajectory. Leverages LLMs’ non-determinism - but hinges on a good eval mechanism (we use an LLM-as-a-Judge). Can increase accuracy without linearly increasing latency - but parallel runs can add up $$.

Source: AI21 Labs (X) | 2026-05-13

Loading related sources…