刚刚,DeepSeek V4 系列更新,架构没变,Agent 能力为何大涨
TL;DR - DeepSeek released V4-Flash-0731 in API public beta, improving official agent benchmarks through renewed post-training without changing model architecture or parameter count. The update also adds native Responses API and Codex workflow support, though independent validation remains limited.
- Official scores reached 82.7 on Terminal Bench 2.1, 54.4 on DeepSWE, and 68.7 on DSBench-FullStack.
- Gains target multi-step software-engineering tasks involving repository inspection, terminal use, code modification, and iterative testing.
- Results depend partly on an unreleased DeepSeek Harness, high reasoning budgets, and internal benchmarks, limiting reproducibility.
- Responses API support simplifies tool calls and agent integration but does not itself increase model intelligence.