从视觉压缩,到视觉记忆:MonkeyOCRv2以重建保留文档页面证据
TL;DR - MonkeyOCRv2 trains document visual encoders with pixel-level reconstruction so compressed tokens retain fine-grained page evidence. This improves OCR fidelity and transfers across seven document-understanding tasks.
- With Qwen3-1.7B and training settings fixed, MonkeyOCRv2-B scored 57.2 across eight benchmarks—13.2 points above OpenVision-B.
- Reconstruction raised shuffled-text accuracy at low resolution from 55.4% to 72.1%, indicating greater reliance on visual evidence rather than language priors.
- The approach improved recognition, formula parsing, text detection, tamper detection, overlapping-text segmentation, document parsing, and understanding.
- The project also introduces MonkeyDoc v2, comprising 113 million document images/elements across 17 languages.