Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Ranking
Overall
78
Content
95
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - CHIVE is an agentic pipeline that evaluates explanations of real-world LLM behavior by testing whether they predict responses to counterfactual prompt edits. Existing interpretability methods showed no predictive uplift, while training on CHIVE experiments improved out-of-distribution generalization.
- Uses counterfactual simulatability as a practical measure of explanation quality.
- Generates thousands of explanations paired with supporting counterfactual evidence.
- Finds no benefit from the evaluated interpretability techniques for predicting counterfactual behavior.
- CHIVE-generated training data generalizes across multiple out-of-distribution settings.
Sources (1)
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - CHIVE is an agentic pipeline that evaluates explanations of real-world LLM behavior by testing whether they predict responses to counterfactual prompt edits. Existing interpretability methods showed no predictive uplift, while training on CHIVE experiments improved out-of-distribution generalization.
- Uses counterfactual simulatability as a practical measure of explanation quality.
- Generates thousands of explanations paired with supporting counterfactual evidence.
- Finds no benefit from the evaluated interpretability techniques for predicting counterfactual behavior.
- CHIVE-generated training data generalizes across multiple out-of-distribution settings.