🛰️ Daily AI Frontier
‹ back to 2026-08-20

DeepSeek V4 Flash不换模型,只靠「自验证」反超Fable 5

WeChat: PaperWeekly LLM Agents 2026-08-19
Representative image for DeepSeek V4 Flash不换模型,只靠「自验证」反超Fable 5

TL;DR - DeepSeek V4 Flash paired with LLM-as-a-Verifier reached 88.0% success on Terminal-Bench 2.1 by generating and self-ranking five agent trajectories. The system reportedly surpassed Claude Fable 5 at roughly $0.11 per task versus $1.30, though the configurations were not directly equivalent.

  • Best-of-3 raised success from 79.4% to 86.5%; Best-of-5 raised it from 78.7% to 88.0%.
  • The same DeepSeek model generated and verified candidates, using token-probability-weighted continuous scores rather than a separately trained verifier.
  • The five-candidate oracle success rate was 96.6%, indicating that better candidate selection remains a major opportunity.
  • Prefix-caching improvements increased the validation-stage cache hit rate from 5.2% to 78.4%, reducing input costs.

view merged work →