DeepSeek V4 Flash不换模型,只靠「自验证」反超Fable 5
Ranking
Overall
85
Content
95
Popularity
60
Observed public metrics from 1 member.
Merged summary
TL;DR - DeepSeek V4 Flash paired with LLM-as-a-Verifier reached 88.0% success on Terminal-Bench 2.1 by generating and self-ranking five agent trajectories. The system reportedly surpassed Claude Fable 5 at roughly $0.11 per task versus $1.30, though the configurations were not directly equivalent.
- Best-of-3 raised success from 79.4% to 86.5%; Best-of-5 raised it from 78.7% to 88.0%.
- The same DeepSeek model generated and verified candidates, using token-probability-weighted continuous scores rather than a separately trained verifier.
- The five-candidate oracle success rate was 96.6%, indicating that better candidate selection remains a major opportunity.
- Prefix-caching improvements increased the validation-stage cache hit rate from 5.2% to 78.4%, reducing input costs.
Sources (1)
DeepSeek V4 Flash不换模型,只靠「自验证」反超Fable 5
Public signals
Hugging Face upvotes 17
TL;DR - DeepSeek V4 Flash paired with LLM-as-a-Verifier reached 88.0% success on Terminal-Bench 2.1 by generating and self-ranking five agent trajectories. The system reportedly surpassed Claude Fable 5 at roughly $0.11 per task versus $1.30, though the configurations were not directly equivalent.
- Best-of-3 raised success from 79.4% to 86.5%; Best-of-5 raised it from 78.7% to 88.0%.
- The same DeepSeek model generated and verified candidates, using token-probability-weighted continuous scores rather than a separately trained verifier.
- The five-candidate oracle success rate was 96.6%, indicating that better candidate selection remains a major opportunity.
- Prefix-caching improvements increased the validation-stage cache hit rate from 5.2% to 78.4%, reducing input costs.