🛰️ Daily AI Frontier
‹ back to 2026-09-23

FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents

Research LLM Agents

Ranking

Overall 79
Content 100
Popularity 29

Observed public metrics from 1 member.

Merged summary

TL;DR - FIRE applies failure-informed runtime instructions and action denials to language-model agents without modifying model weights or user prompts. It substantially improves repeatable task completion, suggesting that harness-level policies can turn existing capabilities into more dependable—and potentially cheaper—agent performance.

  • Across 87 Terminal-Bench 2.1 tasks, FIRE raised pass² from 50.6% to 54.0% for Luna, 55.2% to 60.9% for Terra, and 64.4% to 73.6% for Sol.
  • Sol’s best-of-two success improved by only 1.2 points while pass² rose by 9.2 points, indicating greater consistency rather than major new capability.
  • On 14 tasks, policy-guided Terra achieved 71.4% versus 64.3% for unassisted Sol at about half the cost.
  • In a randomized five-arm experiment, real policies reached 61% success on eligible tasks, compared with 36–43% for no-policy, sham, verification, and reconsideration controls.

Sources (1)

FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents

arXiv cs.AI Nikita Agarwal, Nivedit Jain 2026-09-22 arXiv:2609.26048
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-26 14:14:32.072632 UTC

TL;DR - FIRE applies failure-informed runtime instructions and action denials to language-model agents without modifying model weights or user prompts. It substantially improves repeatable task completion, suggesting that harness-level policies can turn existing capabilities into more dependable—and potentially cheaper—agent performance.

  • Across 87 Terminal-Bench 2.1 tasks, FIRE raised pass² from 50.6% to 54.0% for Luna, 55.2% to 60.9% for Terra, and 64.4% to 73.6% for Sol.
  • Sol’s best-of-two success improved by only 1.2 points while pass² rose by 9.2 points, indicating greater consistency rather than major new capability.
  • On 14 tasks, policy-guided Terra achieved 71.4% versus 64.3% for unassisted Sol at about half the cost.
  • In a randomized five-arm experiment, real policies reached 61% success on eligible tasks, compared with 36–43% for no-policy, sham, verification, and reconsideration controls.
item →