🛰️ Daily AI Frontier
‹ back to 2026-08-18

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

arXiv cs.LG LLM Interpretability Adam Karvonen, Euan Ong, Subhash Kantamneni, Samuel Marks 2026-08-17
Representative image for Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

TL;DR - CHIVE is an agentic pipeline that evaluates explanations of real-world LLM behavior by testing whether they predict responses to counterfactual prompt edits. Existing interpretability methods showed no predictive uplift, while training on CHIVE experiments improved out-of-distribution generalization.

  • Uses counterfactual simulatability as a practical measure of explanation quality.
  • Generates thousands of explanations paired with supporting counterfactual evidence.
  • Finds no benefit from the evaluated interpretability techniques for predicting counterfactual behavior.
  • CHIVE-generated training data generalizes across multiple out-of-distribution settings.

view merged work →