🛰️ Daily AI Frontier
‹ back to 2026-07-27

从视觉压缩,到视觉记忆:MonkeyOCRv2以重建保留文档页面证据

WeChat: 机器之心 Multimodal & Generative 2026-07-26
Representative image for 从视觉压缩,到视觉记忆:MonkeyOCRv2以重建保留文档页面证据

TL;DR - MonkeyOCRv2 trains document visual encoders with pixel-level reconstruction so compressed tokens retain fine-grained page evidence. This improves OCR fidelity and transfers across seven document-understanding tasks.

  • With Qwen3-1.7B and training settings fixed, MonkeyOCRv2-B scored 57.2 across eight benchmarks—13.2 points above OpenVision-B.
  • Reconstruction raised shuffled-text accuracy at low resolution from 55.4% to 72.1%, indicating greater reliance on visual evidence rather than language priors.
  • The approach improved recognition, formula parsing, text detection, tamper detection, overlapping-text segmentation, document parsing, and understanding.
  • The project also introduces MonkeyDoc v2, comprising 113 million document images/elements across 17 languages.

view merged work →