🛰️ Daily AI Frontier
‹ back to 2026-08-13

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

Research LLM Agents

Ranking

Overall 88
Content 100
Popularity 58

Observed public metrics from 1 member.

Merged summary

TL;DR - Harness-IF is a benchmark that evaluates whether coding agents genuinely follow operational rules across five instruction surfaces rather than coincidentally matching default behavior. Results show conventional aggregate scores systematically overstate compliance.

  • The benchmark includes 60 multi-turn coding tasks and 256 evaluated rules drawn from a 642-rule library.
  • Across 12 frontier models, overall accuracy was 72.1–85.9%, while Against-Prior Accuracy was lower at 66.1–78.6%.
  • Every model performed worse on rules opposing its defaults, with gaps of 3.6–7.4 percentage points.
  • A conflict pilot found instruction precedence did not follow prompt depth; system prompts, project files, and user instructions outranked tool and skill descriptions.

Sources (1)

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

arXiv cs.AI Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen, Haodong Wang, Zhongfei Hou, Changxin Pu, Shen Yan, Wenhao Huang 2026-08-12 arXiv:2608.11727
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-08-26 14:32:30.232563 UTC

TL;DR - Harness-IF is a benchmark that evaluates whether coding agents genuinely follow operational rules across five instruction surfaces rather than coincidentally matching default behavior. Results show conventional aggregate scores systematically overstate compliance.

  • The benchmark includes 60 multi-turn coding tasks and 256 evaluated rules drawn from a 642-rule library.
  • Across 12 frontier models, overall accuracy was 72.1–85.9%, while Against-Prior Accuracy was lower at 66.1–78.6%.
  • Every model performed worse on rules opposing its defaults, with gaps of 3.6–7.4 percentage points.
  • A conflict pilot found instruction precedence did not follow prompt depth; system prompts, project files, and user instructions outranked tool and skill descriptions.
item →