Untangling the Mechanisms of Misleading Context in Medical Question Answering
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - This study examines how fabricated evidence and unsupported assertions corrupt medical reasoning models. Models were especially vulnerable to bare assertions, while access to open reasoning traces substantially improved detection of corrupted decisions.
- Across three reasoning models, asserted answers were adopted 10–27 percentage points more often than answers suggested by fabricated evidence.
- Misleading cues appeared in 81–98% of reasoning traces but only 7–90% of final responses, with assertions disclosed less often than evidence-based cues.
- Fabricated evidence influenced reasoning early and accumulated, whereas bare assertions tended to redirect the conclusion near the end.
- With guidance, an LLM monitor detected 78% of corrupted decisions at a 5% false-positive rate from an open model’s trace, versus at most 32% from final responses.
Sources (1)
Untangling the Mechanisms of Misleading Context in Medical Question Answering
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This study examines how fabricated evidence and unsupported assertions corrupt medical reasoning models. Models were especially vulnerable to bare assertions, while access to open reasoning traces substantially improved detection of corrupted decisions.
- Across three reasoning models, asserted answers were adopted 10–27 percentage points more often than answers suggested by fabricated evidence.
- Misleading cues appeared in 81–98% of reasoning traces but only 7–90% of final responses, with assertions disclosed less often than evidence-based cues.
- Fabricated evidence influenced reasoning early and accumulated, whereas bare assertions tended to redirect the conclusion near the end.
- With guidance, an LLM monitor detected 78% of corrupted decisions at a 5% false-positive rate from an open model’s trace, versus at most 32% from final responses.