Model Releases

4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filte…

4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filters failing patches, pushing our final result to 60.9% - surp

DGX agentx-post
model-releasesai21-labs--x

4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filters failing patches, pushing our final result to 60.9% - surpassing Claude Code (60.9% vs 56.2%) at the same cost.

Source: AI21 Labs (X) | 2026-06-04

Loading related sources…