🛰️ Daily AI Frontier
‹ back to 2026-07-22

刚刚,超1亿文档图像数据开源:MonkeyOCRv2想给视觉文档补上一块公共底座

Industry & News Multimodal & Generative

Merged summary

TL;DR - MonkeyOCRv2 introduces a document-native visual encoder trained with text generation and pixel reconstruction to preserve character- and layout-level evidence. Its team also released MonkeyDoc v2, a 100+ TB dataset containing 113 million document images across 17 languages.

  • Pixel reconstruction improved shuffled-text recognition at low resolution from 55.4% to 72.1%, reducing reliance on language-model guessing.
  • With a frozen Qwen3-1.7B backend, MonkeyOCRv2-B averaged 57.2 across eight document-understanding benchmarks versus 44.0 for OpenVision-B.
  • A 0.7B document parser built around the frozen encoder scored 83.3 on multilingual MDPBench, 2.8 points above a prior 3B open-source baseline.
  • MonkeyDoc v2 includes 8 million full pages and 105 million cropped elements, combining 61 million real-world and 52 million synthetic samples.

Sources (1)

刚刚,超1亿文档图像数据开源:MonkeyOCRv2想给视觉文档补上一块公共底座

WeChat: 极市平台 2026-07-21 arXiv:2607.11562

TL;DR - MonkeyOCRv2 introduces a document-native visual encoder trained with text generation and pixel reconstruction to preserve character- and layout-level evidence. Its team also released MonkeyDoc v2, a 100+ TB dataset containing 113 million document images across 17 languages.

  • Pixel reconstruction improved shuffled-text recognition at low resolution from 55.4% to 72.1%, reducing reliance on language-model guessing.
  • With a frozen Qwen3-1.7B backend, MonkeyOCRv2-B averaged 57.2 across eight document-understanding benchmarks versus 44.0 for OpenVision-B.
  • A 0.7B document parser built around the frozen encoder scored 83.3 on multilingual MDPBench, 2.8 points above a prior 3B open-source baseline.
  • MonkeyDoc v2 includes 8 million full pages and 105 million cropped elements, combining 61 million real-world and 52 million synthetic samples.
item →