Agents
3/5 An example: In instance psf__requests-1724, the gold fix is 2 lines. Our agent’s functional fix was 8 lines. The LLM judge rejected the …
3/5 An example: In instance psf__requests-1724, the gold fix is 2 lines. Our agent’s functional fix was 8 lines. The LLM judge rejected the correct 8-liner as 'messy' and 'redundant,' choosing a clean
3/5 An example: In instance psf__requests-1724, the gold fix is 2 lines. Our agent’s functional fix was 8 lines. The LLM judge rejected the correct 8-liner as "messy" and "redundant," choosing a clean but non-functional fix instead. 🔗 See full patch in the blog: https://www.ai21.com/blog/gold-like-answers-benchmarks/?utm_source=org-twitter
Related
- 5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of th…
- Reasoning Graphs: Deterministic Agent Accuracy through Evidence-Centric Chain-of-Thought Feedback
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
- 智谱直接把开源 Agent 拉到新高度了! GLM-5.1 正式开源: ✅ 开源模型里 SWE-Bench Pro 拿下 #1(58.4),全球第 3 ✅ 真正长时程 Agent:自主运行 8 小时、几千次迭代 + 自审循环 ✅ 从零搭出一个带 50+ App 的完整 Linux…
Source: AI21 Labs (X) | 2026-04-15