HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
arXiv:2607.25398v1 Announce Type: new Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context,