🛰️ Daily AI Frontier
‹ back to 2026-08-13

连夜实测DeepSeek V4 Pro 正式版,低于预期,不推荐接入Codex

Opinions LLMs & Foundation Models

Ranking

Overall 71
Content 80
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 连夜实测DeepSeek V4 Pro 正式版,低于预期,不推荐接入Codex

Merged summary

TL;DR - A hands-on review finds DeepSeek V4 Pro inexpensive and operationally robust, but disappointing as a primary Codex model due to limited progress visibility, aggressive reuse of prior context, and weaker-than-expected frontend and writing performance.

  • A 41.39M-token API test achieved 100% request success, 52.3K tok/s aggregate throughput, and roughly 96% cache hits for ¥8.5.
  • Reasoning comprised 84.5% of output tokens, while sparse user-facing updates made long-running tasks difficult to monitor.
  • The model repeatedly reused artifacts and conclusions from earlier sessions, improving efficiency but raising concerns about stale or inappropriate context reuse.
  • Latency reached 4.9 seconds at p50 and 12.3 seconds at p99 under the reported workload; frontend generation and writing also trailed expectations.

Sources (1)

连夜实测DeepSeek V4 Pro 正式版,低于预期,不推荐接入Codex

WeChat: 夕小瑶科技说 2026-08-12
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-12 14:27:59.839564 UTC

TL;DR - A hands-on review finds DeepSeek V4 Pro inexpensive and operationally robust, but disappointing as a primary Codex model due to limited progress visibility, aggressive reuse of prior context, and weaker-than-expected frontend and writing performance.

  • A 41.39M-token API test achieved 100% request success, 52.3K tok/s aggregate throughput, and roughly 96% cache hits for ¥8.5.
  • Reasoning comprised 84.5% of output tokens, while sparse user-facing updates made long-running tasks difficult to monitor.
  • The model repeatedly reused artifacts and conclusions from earlier sessions, improving efficiency but raising concerns about stale or inappropriate context reuse.
  • Latency reached 4.9 seconds at p50 and 12.3 seconds at p99 under the reported workload; frontend generation and writing also trailed expectations.
item →