🛰️ Daily AI Frontier
‹ back to 2026-08-07

ICML 2026 | 多模态思维链真的靠谱吗?川大新框架拒绝噪声思维干扰

Research Multimodal & Generative

Ranking

Overall 74
Content 80
Popularity 59

Observed public metrics from 1 member.

Representative image for ICML 2026 | 多模态思维链真的靠谱吗?川大新框架拒绝噪声思维干扰

Merged summary

TL;DR - An ICML 2026 paper from Sichuan University, "Reliable Thinking with Images" (RTWI), tackles "noisy thoughts" in multimodal chain-of-thought reasoning, where wrong visual cues or faulty reasoning cascade into wrong answers. It offers a plug-and-play test-time scaling framework that raises accuracy while cutting inference cost.

  • Frames the Thinking-with-Images (TWI) pipeline as two stages — cue mining (tool-driven visual clue extraction) and answer reasoning — and shows via error analysis that incorrect answers usually trace back to errors in one of these stages.
  • Instead of modeling continuous visual uncertainty directly, RTWI uses a text-centric reliability measure based on token entropy per stage, since the tool-calling text instruction is the precondition for obtaining correct clues.
  • Two empirical observations drive the method: reliability correlation (higher stage reliability ↔ higher answer accuracy) and reliability jump (correct visual clues yield larger reliability gains from cue mining to answer reasoning).
  • The framework combines dual-stage filtering (percentile thresholds discard unreliable paths) with reliability-weighted voting replacing majority voting; tested on Qwen3-VL and DeepEyes across high-resolution, TWI-specific, multimodal math, and open-ended VQA benchmarks, reporting accuracy gains, competitive token-saving rates via online early stopping, and consistent benefits across model scales.

Sources (1)

ICML 2026 | 多模态思维链真的靠谱吗?川大新框架拒绝噪声思维干扰

WeChat: PaperWeekly 2026-08-06 arXiv:2602.12916
Public signals Semantic Scholar citations 4 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 4 · Influential citations 0 X · N/A Fetched 2026-08-31 14:25:47.286593 UTC

TL;DR - An ICML 2026 paper from Sichuan University, "Reliable Thinking with Images" (RTWI), tackles "noisy thoughts" in multimodal chain-of-thought reasoning, where wrong visual cues or faulty reasoning cascade into wrong answers. It offers a plug-and-play test-time scaling framework that raises accuracy while cutting inference cost.

  • Frames the Thinking-with-Images (TWI) pipeline as two stages — cue mining (tool-driven visual clue extraction) and answer reasoning — and shows via error analysis that incorrect answers usually trace back to errors in one of these stages.
  • Instead of modeling continuous visual uncertainty directly, RTWI uses a text-centric reliability measure based on token entropy per stage, since the tool-calling text instruction is the precondition for obtaining correct clues.
  • Two empirical observations drive the method: reliability correlation (higher stage reliability ↔ higher answer accuracy) and reliability jump (correct visual clues yield larger reliability gains from cue mining to answer reasoning).
  • The framework combines dual-stage filtering (percentile thresholds discard unreliable paths) with reliability-weighted voting replacing majority voting; tested on Qwen3-VL and DeepEyes across high-resolution, TWI-specific, multimodal math, and open-ended VQA benchmarks, reporting accuracy gains, competitive token-saving rates via online early stopping, and consistent benefits across model scales.
item →