Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
TL;DR - CHIVE is an agentic pipeline that evaluates explanations of real-world LLM behavior by testing whether they predict responses to counterfactual prompt edits. Existing interpretability methods showed no predictive uplift, while training on CHIVE experiments improved out-of-distribution generalization.
- Uses counterfactual simulatability as a practical measure of explanation quality.
- Generates thousands of explanations paired with supporting counterfactual evidence.
- Finds no benefit from the evaluated interpretability techniques for predicting counterfactual behavior.
- CHIVE-generated training data generalizes across multiple out-of-distribution settings.