🛰️ Daily AI Frontier
‹ back to 2026-08-02

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

Research Multimodal & Generative

Ranking

Overall 75
Content 90
Popularity 40

Observed public metrics from 1 member.

Representative image for FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

Merged summary

TL;DR - FaithEyes is a multi-agent framework that trains vision-language models to use image-processing tools only when their outputs genuinely aid reasoning. It improves tool faithfulness while maintaining competitive or superior benchmark accuracy.

  • A VLM judges whether each cropped or manipulated process image helps answer the question.
  • Helpfulness judgments guide subsequent reasoning and scale rewards to discourage decorative or misaligned tool calls.
  • At inference, the model acts as its own judging subagent, avoiding reliance on an external evaluator.
  • Training uses a two-stage supervised fine-tuning and reinforcement learning pipeline on adapted open-source data.

Sources (1)

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

arXiv cs.CV Haoqing Wang, Xingrun Xing, Wei Xia, Ziheng Li, Yehui Tang 2026-07-30 arXiv:2607.28225
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-06 02:06:57.135395 UTC

TL;DR - FaithEyes is a multi-agent framework that trains vision-language models to use image-processing tools only when their outputs genuinely aid reasoning. It improves tool faithfulness while maintaining competitive or superior benchmark accuracy.

  • A VLM judges whether each cropped or manipulated process image helps answer the question.
  • Helpfulness judgments guide subsequent reasoning and scale rewards to discourage decorative or misaligned tool calls.
  • At inference, the model acts as its own judging subagent, avoiding reliance on an external evaluator.
  • Training uses a two-stage supervised fine-tuning and reinforcement learning pipeline on adapted open-source data.
item →