🛰️ Daily AI Frontier
‹ back to 2026-09-23

Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding

Research Efficiency & Systems

Ranking

Overall 75
Content 95
Popularity 29

Observed public metrics from 1 member.

Representative image for Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding

Merged summary

TL;DR - GravityOCR combines diffusion-based parallel drafting with autoregressive verification in one parameter-shared OCR model. It accelerates document decoding while retaining accuracy close to the original GLM-OCR model.

  • The model drafts multiple tokens in parallel, then causally verifies them before commitment to reduce errors.
  • Its autoregressive path supports GRPO using sequence- and structure-level OCR rewards without diffusion-trajectory likelihood estimation.
  • On OmniDocBench v1.6, GRPO improved the Overall score from 94.92 to 95.16; the final model scored near GLM-OCR’s 95.48.
  • In SGLang, GravityOCR committed 9.7 tokens per forward pass on average, delivering 3.94Ă— decode-only and 1.32Ă— end-to-end speedups over autoregressive decoding.

Sources (1)

Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding

arXiv cs.CL Dohyun Kim, Sungjun Han, Hyungguk Kim, Yusik Kim, Jamin Shin, Paul Hongsuck Seo, Hongjoon Ahn 2026-09-22 arXiv:2609.26638
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-24 14:15:52.205318 UTC

TL;DR - GravityOCR combines diffusion-based parallel drafting with autoregressive verification in one parameter-shared OCR model. It accelerates document decoding while retaining accuracy close to the original GLM-OCR model.

  • The model drafts multiple tokens in parallel, then causally verifies them before commitment to reduce errors.
  • Its autoregressive path supports GRPO using sequence- and structure-level OCR rewards without diffusion-trajectory likelihood estimation.
  • On OmniDocBench v1.6, GRPO improved the Overall score from 94.92 to 95.16; the final model scored near GLM-OCR’s 95.48.
  • In SGLang, GravityOCR committed 9.7 tokens per forward pass on average, delivering 3.94Ă— decode-only and 1.32Ă— end-to-end speedups over autoregressive decoding.
item →