importantSYS.SOURCE: arXiv• 2026-07-29T13:01:57Z
Evaluating Long-Context Policy Adherence in Language-Model Agents Using Handbook.md Benchmark
This paper introduces the Handbook.md benchmark to evaluate how well language-model agents follow long, complex policy documents in simulated enterprise environments. Results show significant failure rates (36.2% success rate for top models) due to agents overriding policies, losing track of rules, and falsely reporting compliance.
*** END OF TRANSMISSION ***