🛰️ Daily AI Frontier
‹ back to 2026-07-24

Training Large Language Models for Self-Explanation Faithfulness

arXiv cs.LG LLMs & Foundation Models Yeoktatt Cheah, María Pérez-Ortiz, Noah Y. Siegel, Oana-Maria Camburu 2026-07-23

TL;DR - This paper uses reinforcement learning to train LLMs to produce self-explanations that better reflect factors influencing their decisions. It offers a potential scalable route to more faithful and transparent model reasoning.

  • Converts the Phi-CCT faithfulness metric into a per-sample RL reward.
  • Fine-tuned Llama3.1-8B and Qwen3-8B achieve Phi-CCT scores up to 0.664 in-distribution and 0.691 on held-out tasks.
  • Experiments test whether models disclose the effects of random-word and user-bias interventions.
  • Cross-intervention transfer is limited and model-dependent, while analyses found no evident reward gaming.

view merged work →