Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress
TL;DR - This paper proves that prediction-based safeguards such as accuracy, calibration, and conformal coverage cannot by themselves certify trustworthy AI. It proposes a “competence envelope” that combines prediction and explanation certification to expose otherwise invisible failures.
- A reliable model and a compromised model can satisfy identical prediction-side certificates while differing arbitrarily in explanation fidelity and deployment behavior.
- Detecting this separation requires evidence about the model’s decision mechanism, not merely its outputs.
- The competence envelope provides a deployable criterion integrating prediction performance with explanation fidelity.
- Experiments across multiple datasets and model classes reveal failure modes missed by prediction-only certification.