Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
Ranking
Overall
88
Content
100
Popularity
58
Observed public metrics from 1 member.
Merged summary
TL;DR - Harness-IF is a benchmark that evaluates whether coding agents genuinely follow operational rules across five instruction surfaces rather than coincidentally matching default behavior. Results show conventional aggregate scores systematically overstate compliance.
- The benchmark includes 60 multi-turn coding tasks and 256 evaluated rules drawn from a 642-rule library.
- Across 12 frontier models, overall accuracy was 72.1–85.9%, while Against-Prior Accuracy was lower at 66.1–78.6%.
- Every model performed worse on rules opposing its defaults, with gaps of 3.6–7.4 percentage points.
- A conflict pilot found instruction precedence did not follow prompt depth; system prompts, project files, and user instructions outranked tool and skill descriptions.
Sources (1)
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
Public signals
Semantic Scholar citations 1 · Semantic Scholar influential citations 0
TL;DR - Harness-IF is a benchmark that evaluates whether coding agents genuinely follow operational rules across five instruction surfaces rather than coincidentally matching default behavior. Results show conventional aggregate scores systematically overstate compliance.
- The benchmark includes 60 multi-turn coding tasks and 256 evaluated rules drawn from a 642-rule library.
- Across 12 frontier models, overall accuracy was 72.1–85.9%, while Against-Prior Accuracy was lower at 66.1–78.6%.
- Every model performed worse on rules opposing its defaults, with gaps of 3.6–7.4 percentage points.
- A conflict pilot found instruction precedence did not follow prompt depth; system prompts, project files, and user instructions outranked tool and skill descriptions.