Model Releases
If you maintain an AGENTS.md or a CLAUDE.md, this one is worth your time. (bookmark it) Researchers traced 94K development events across 557…
If you maintain an AGENTS.md or a CLAUDE.md, this one is worth your time. (bookmark it) Researchers traced 94K development events across 557 agentic coding sessions, plus 690K file-level change record
If you maintain an AGENTS.md or a CLAUDE.md, this one is worth your time. (bookmark it) Researchers traced 94K development events across 557 agentic coding sessions, plus 690K file-level change records from 33K agentic pull requests. Instruction files and working notes account for 60.5% of everything agents read. Classical technical docs get 10.6%. API references get 1.3%. Reading docs is associated with less immediate testing, at an adjusted odds ratio of 0.39. And consultation is self-initiated 70.2% of the time, against 7.5% driven by a failure. In multi-commit agentic pull requests, code gets touched first 4.7x more often. Paper: https://arxiv.org/abs/2608.20195 Track more trending AI papers in our academy: https://academy.dair.ai/
Related
- Why does CLAUDE.MD keep growing? If you maintain a CLAUDE.md or an AGENTS.md, this one is worth your time. (bookmark it) This work traces wh…
- If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, …
- New research with Microsoft's and colleagues on training agents inside the harnesses they actually run in. (bookmark it) Why it matters: Age…
- Great paper on self-improving agent harnesses. (bookmark it) If you maintain a production agent harness, finding every file behind one behav…
Source: DAIR.AI (X) | 2026-08-23