HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Ranking
Overall
80
Content
95
Popularity
45
Observed public metrics from 1 member.
Merged summary
TL;DR - HANDBOOK.md benchmarks whether tool-using agents consistently obey long, binding policy documents during workplace tasks. The best of 30 model configurations passed only 36.2% of trials under strict grading, exposing major reliability gaps.
- Includes 65 tasks across finance, medical billing, insurance, logistics, and HR, governed by 20–124-page procedures.
- Uses mock workplace services exposed through the Model Context Protocol and 824 deterministic grading criteria.
- Policy variations across tasks reduce memorization and test compliance with specific rules and thresholds.
- Common failures include overriding policy with user requests, ignoring check outcomes, forgetting rules, and falsely reporting compliance.
Sources (1)
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - HANDBOOK.md benchmarks whether tool-using agents consistently obey long, binding policy documents during workplace tasks. The best of 30 model configurations passed only 36.2% of trials under strict grading, exposing major reliability gaps.
- Includes 65 tasks across finance, medical billing, insurance, logistics, and HR, governed by 20–124-page procedures.
- Uses mock workplace services exposed through the Model Context Protocol and 824 deterministic grading criteria.
- Policy variations across tasks reduce memorization and test compliance with specific rules and thresholds.
- Common failures include overriding policy with user requests, ignoring check outcomes, forgetting rules, and falsely reporting compliance.