🛰️ Daily AI Frontier
‹ back to 2026-08-13

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

arXiv cs.AI LLM Agents Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen, Haodong Wang, Zhongfei Hou, Changxin Pu, Shen Yan, Wenhao Huang 2026-08-12

TL;DR - Harness-IF is a benchmark that evaluates whether coding agents genuinely follow operational rules across five instruction surfaces rather than coincidentally matching default behavior. Results show conventional aggregate scores systematically overstate compliance.

  • The benchmark includes 60 multi-turn coding tasks and 256 evaluated rules drawn from a 642-rule library.
  • Across 12 frontier models, overall accuracy was 72.1–85.9%, while Against-Prior Accuracy was lower at 66.1–78.6%.
  • Every model performed worse on rules opposing its defaults, with gaps of 3.6–7.4 percentage points.
  • A conflict pilot found instruction precedence did not follow prompt depth; system prompts, project files, and user instructions outranked tool and skill descriptions.

view merged work →