Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
TL;DR - Harness-IF is a benchmark that evaluates whether coding agents genuinely follow operational rules across five instruction surfaces rather than coincidentally matching default behavior. Results show conventional aggregate scores systematically overstate compliance.
- The benchmark includes 60 multi-turn coding tasks and 256 evaluated rules drawn from a 642-rule library.
- Across 12 frontier models, overall accuracy was 72.1–85.9%, while Against-Prior Accuracy was lower at 66.1–78.6%.
- Every model performed worse on rules opposing its defaults, with gaps of 3.6–7.4 percentage points.
- A conflict pilot found instruction precedence did not follow prompt depth; system prompts, project files, and user instructions outranked tool and skill descriptions.