🛰️ Daily AI Frontier
‹ back to 2026-09-18

SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment

arXiv cs.CL LLMs & Foundation Models Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu 2026-09-17

TL;DR - SAFARI is an industrial benchmark for evaluating LLM-assisted automotive hazard analysis and risk assessment under ISO 26262. Tests show that frontier LLMs can generate plausible hazard narratives but remain unreliable at standards-based risk classification.

  • Includes 3,000 de-identified industrial HARA cases covering open-ended hazard analysis and standards-grounded risk assessment.
  • Introduces a reference-anchored LLM-as-a-judge protocol that correlates strongly with expert evaluation.
  • Across nine frontier LLMs, the best ASIL classification macro-F1 was only 0.261.
  • Chain-of-Thought prompting offered limited gains and often worsened categorical assessment; key errors involved missing scenario context and misjudging controllability.

view merged work →