🛰️ Daily AI Frontier
‹ back to 2026-09-26

Who Holds the Pen? Let Specifications, Not Agents, Sign Off

arXiv cs.AI LLM Agents Haiqing Li, Xin Ma, Yinhao Wu, Wenliang Zhong, Feng Jiang, Thao M. Dang, Xiao Hu, Hehuan Ma, Yuzhi Guo, Junzhou Huang 2026-09-24
Representative image for Who Holds the Pen? Let Specifications, Not Agents, Sign Off

TL;DR - This paper argues that LLM agents should not certify their own task completion; instead, external specifications and admissible evidence should determine whether requirements are satisfied. Its SpecHarness framework turns visible specifications into tracked obligations that govern execution and finalization.

  • Across seven models on SkillsBench, only 79.6%–86.4% of 509 source-grounded task directions were satisfied.
  • Agent completion claims exceeded official evaluator pass rates by 28.7–37.9 percentage points, highlighting a substantial state–authority gap.
  • SpecHarness compiles specifications into source-linked, versioned obligations and validates verifiable requirements during runtime.
  • Ambiguous or subjective requirements remain advisory, while qualified evidence providers establish authoritative state for checkable requirements.

view merged work →