🛰️ Daily AI Frontier
‹ back to 2026-08-28

SWE-Prime: Fewer Trajectories, Better Performance

Research LLM Agents

Ranking

Overall 85
Content 95
Popularity 61

Observed public metrics from 1 member.

Representative image for SWE-Prime: Fewer Trajectories, Better Performance

Merged summary

TL;DR - SWE-Prime is a two-stage data-selection method for supervised fine-tuning of software-engineering agents that filters successful trajectories by quality and representativeness, then selects useful segments for loss computation. Using only 10% of trajectories outperformed training on the full resolved dataset, showing that cleaner supervision can beat greater data volume.

  • Screens trajectories using process quality, result quality, and dataset representativeness.
  • Evaluates semantic step segments for solution contribution, learnability, and potential risks.
  • Retains all segments as context during training but computes loss only on selected segments.
  • Achieved relative gains of up to 12.2% on SWE-Bench Pro and 24.2% on SWE-Bench Verified.

Sources (1)

SWE-Prime: Fewer Trajectories, Better Performance

arXiv cs.SE Dewu Zheng, Ruizhe Ye, Yanlin Wang, Yang Ye, Hongyu Zhang, Ensheng Shi, Xilin Liu, Yuchi Ma, Jianxing Yu, Zibin Zheng 2026-08-27 arXiv:2608.27449
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-09-19 14:21:26.531035 UTC

TL;DR - SWE-Prime is a two-stage data-selection method for supervised fine-tuning of software-engineering agents that filters successful trajectories by quality and representativeness, then selects useful segments for loss computation. Using only 10% of trajectories outperformed training on the full resolved dataset, showing that cleaner supervision can beat greater data volume.

  • Screens trajectories using process quality, result quality, and dataset representativeness.
  • Evaluates semantic step segments for solution contribution, learnability, and potential risks.
  • Retains all segments as context during training but computes loss only on selected segments.
  • Achieved relative gains of up to 12.2% on SWE-Bench Pro and 24.2% on SWE-Bench Verified.
item →