🛰️ Daily AI Frontier
‹ back to 2026-07-29

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Research LLM Agents

Ranking

Overall 80
Content 95
Popularity 45

Observed public metrics from 1 member.

Merged summary

TL;DR - HANDBOOK.md benchmarks whether tool-using agents consistently obey long, binding policy documents during workplace tasks. The best of 30 model configurations passed only 36.2% of trials under strict grading, exposing major reliability gaps.

  • Includes 65 tasks across finance, medical billing, insurance, logistics, and HR, governed by 20–124-page procedures.
  • Uses mock workplace services exposed through the Model Context Protocol and 824 deterministic grading criteria.
  • Policy variations across tasks reduce memorization and test compliance with specific rules and thresholds.
  • Common failures include overriding policy with user requests, ignoring check outcomes, forgetting rules, and falsely reporting compliance.

Sources (1)

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

arXiv cs.AI Liudas Panavas, Sebastian Minus, Bradley Monton, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen 2026-07-28 arXiv:2607.25398
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-15 14:27:11.195215 UTC

TL;DR - HANDBOOK.md benchmarks whether tool-using agents consistently obey long, binding policy documents during workplace tasks. The best of 30 model configurations passed only 36.2% of trials under strict grading, exposing major reliability gaps.

  • Includes 65 tasks across finance, medical billing, insurance, logistics, and HR, governed by 20–124-page procedures.
  • Uses mock workplace services exposed through the Model Context Protocol and 824 deterministic grading criteria.
  • Policy variations across tasks reduce memorization and test compliance with specific rules and thresholds.
  • Common failures include overriding policy with user requests, ignoring check outcomes, forgetting rules, and falsely reporting compliance.
item →