Agents

Highly-recommended read. Aligns with what I see in my own harness: > Pi harness got the same success rate as harnesses from the LLM vendors …

Highly-recommended read. Aligns with what I see in my own harness: > Pi harness got the same success rate as harnesses from the LLM vendors with Opus and GPT, but at 2x less cost > GLM 5.2 was a major

DGX agentx-post
agentsdair-ai--x

Highly-recommended read. Aligns with what I see in my own harness: > Pi harness got the same success rate as harnesses from the LLM vendors with Opus and GPT, but at 2x less cost > GLM 5.2 was a major step forward in open-source coding agent performance The harness bloat/rot is real! We benchmarked coding agents on our own internal tasks at Databricks and learned a lot! There are many surprising opportunities to lower cost and increase quality, and many models including open source ones are truly competitive now. 🧵

Source: DAIR.AI (X) | 2026-07-08

Loading related sources…