🛰️ Daily AI Frontier
‹ back to 2026-09-18

SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment

Research LLMs & Foundation Models

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Merged summary

TL;DR - SAFARI is an industrial benchmark for evaluating LLM-assisted automotive hazard analysis and risk assessment under ISO 26262. Tests show that frontier LLMs can generate plausible hazard narratives but remain unreliable at standards-based risk classification.

  • Includes 3,000 de-identified industrial HARA cases covering open-ended hazard analysis and standards-grounded risk assessment.
  • Introduces a reference-anchored LLM-as-a-judge protocol that correlates strongly with expert evaluation.
  • Across nine frontier LLMs, the best ASIL classification macro-F1 was only 0.261.
  • Chain-of-Thought prompting offered limited gains and often worsened categorical assessment; key errors involved missing scenario context and misjudging controllability.

Sources (1)

SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment

arXiv cs.CL Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu 2026-09-17 arXiv:2609.20584
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:15:27.683961 UTC

TL;DR - SAFARI is an industrial benchmark for evaluating LLM-assisted automotive hazard analysis and risk assessment under ISO 26262. Tests show that frontier LLMs can generate plausible hazard narratives but remain unreliable at standards-based risk classification.

  • Includes 3,000 de-identified industrial HARA cases covering open-ended hazard analysis and standards-grounded risk assessment.
  • Introduces a reference-anchored LLM-as-a-judge protocol that correlates strongly with expert evaluation.
  • Across nine frontier LLMs, the best ASIL classification macro-F1 was only 0.261.
  • Chain-of-Thought prompting offered limited gains and often worsened categorical assessment; key errors involved missing scenario context and misjudging controllability.
item →