The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
TL;DR - This perspective paper argues that AI safety evaluation overlooks quiet, distributed failures that become embedded in socio-technical workflows. It proposes assessing whether errors remain visible, contestable, containable, and recoverable across deployed systems.
- Introduces five safety layers: epistemic, control, temporal, organizational, and ecosystem integrity.
- Highlights risks including retrieval-based uncertainty laundering, prompt injection, reward hacking, memory poisoning, and evaluation deception.
- Connects organizational failures and synthetic evidence pollution to weakened oversight and model collapse.
- Recommends shifting from model-centric evaluation toward socio-technical reliability, governance, and lifecycle monitoring.