🛰️ Daily AI Frontier
‹ back to 2026-09-23

FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents

arXiv cs.AI LLM Agents Nikita Agarwal, Nivedit Jain 2026-09-22

TL;DR - FIRE applies failure-informed runtime instructions and action denials to language-model agents without modifying model weights or user prompts. It substantially improves repeatable task completion, suggesting that harness-level policies can turn existing capabilities into more dependable—and potentially cheaper—agent performance.

  • Across 87 Terminal-Bench 2.1 tasks, FIRE raised pass² from 50.6% to 54.0% for Luna, 55.2% to 60.9% for Terra, and 64.4% to 73.6% for Sol.
  • Sol’s best-of-two success improved by only 1.2 points while pass² rose by 9.2 points, indicating greater consistency rather than major new capability.
  • On 14 tasks, policy-guided Terra achieved 71.4% versus 64.3% for unassisted Sol at about half the cost.
  • In a randomized five-arm experiment, real policies reached 61% success on eligible tasks, compared with 36–43% for no-policy, sham, verification, and reconsideration controls.

view merged work →