🛰️ Daily AI Frontier
‹ back to 2026-08-23

刚刚,DeepSeek首个多模态模型发布!!

WeChat: 极市平台 Multimodal & Generative 2026-08-21
Representative image for 刚刚,DeepSeek首个多模态模型发布!!

TL;DR - DeepSeek released its first experimental vision-language model, deepseek-v4-flash-vision-exp, adding image understanding to V4 Flash at the same API pricing. The launch expands DeepSeek’s low-cost offering toward multimodal agents that can interpret screenshots, charts, documents, and interfaces.

  • Handles up to 600 images per request, supports JPEG, PNG, GIF, and WebP, and accepts Base64 data, public URLs, or Files API references.
  • Supports images up to 8192 pixels on the longest side, with configurable detail levels and a maximum charge of 384 tokens per image.
  • Integrates through OpenAI Chat Completions, Responses API, and Anthropic-compatible formats, easing adoption in existing agent and developer tooling.
  • The model remains experimental; reported agent-benchmark performance near Claude Opus 4.8 awaits fuller documentation and independent evaluation.

view merged work →