Do Reasoning Representations Help Humans Evaluate LLM Outputs?
TL;DR - A controlled human study finds that reasoning formats users prefer are not necessarily the ones that best help them evaluate LLM outputs. Simple chain-of-thought traces outperform planning- and decomposition-based formats for verification, trust calibration, and interpretability.
- The study compares six reasoning formats across tasks of varying complexity using randomized domains, problem instances, and presentation order.
- Participants favored planning- and decomposition-based representations, despite simpler chain-of-thought traces better supporting error detection and evaluation.
- Preferred formats produced calibration risks, including more false alarms on correct traces.
- Participants sometimes reported high trust while remaining unwilling to verify the reasoning themselves.