刚刚,DeepSeek首个多模态模型发布!!
Ranking
Overall
68
Content
75
Popularity
N/A
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - DeepSeek released its first experimental vision-language model, deepseek-v4-flash-vision-exp, adding image understanding to V4 Flash at the same API pricing. The launch expands DeepSeek’s low-cost offering toward multimodal agents that can interpret screenshots, charts, documents, and interfaces.
- Handles up to 600 images per request, supports JPEG, PNG, GIF, and WebP, and accepts Base64 data, public URLs, or Files API references.
- Supports images up to 8192 pixels on the longest side, with configurable detail levels and a maximum charge of 384 tokens per image.
- Integrates through OpenAI Chat Completions, Responses API, and Anthropic-compatible formats, easing adoption in existing agent and developer tooling.
- The model remains experimental; reported agent-benchmark performance near Claude Opus 4.8 awaits fuller documentation and independent evaluation.
Sources (1)
刚刚,DeepSeek首个多模态模型发布!!
Public signals
N/A
TL;DR - DeepSeek released its first experimental vision-language model, deepseek-v4-flash-vision-exp, adding image understanding to V4 Flash at the same API pricing. The launch expands DeepSeek’s low-cost offering toward multimodal agents that can interpret screenshots, charts, documents, and interfaces.
- Handles up to 600 images per request, supports JPEG, PNG, GIF, and WebP, and accepts Base64 data, public URLs, or Files API references.
- Supports images up to 8192 pixels on the longest side, with configurable detail levels and a maximum charge of 384 tokens per image.
- Integrates through OpenAI Chat Completions, Responses API, and Anthropic-compatible formats, easing adoption in existing agent and developer tooling.
- The model remains experimental; reported agent-benchmark performance near Claude Opus 4.8 awaits fuller documentation and independent evaluation.