Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
TL;DR - Aphanta is a diagnostic framework for testing whether image-edited intermediates improve multimodal reasoning. It finds that image editing is useful as a specialized visual workspace for certain tasks, but not as a universal reasoning mechanism.
- Compares direct reasoning, editor-assisted reasoning, and reasoning with idealized reference intermediates to distinguish theoretical headroom from current editor utility.
- Across 20 candidate tasks, gains were concentrated in visual cue injection, grounding, and counterfactual state realization.
- Symbol-sensitive construction and structural extrapolation were substantially less reliable.
- On selected positive tasks, a consolidated Qwen pipeline improved mean score from 0.343 to 0.445, a 10.2-point absolute and 29.7% relative gain.