🛰️ Daily AI Frontier
‹ back to 2026-08-21

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Research LLM Agents

Ranking

Overall 86
Content 90
Popularity 77

Observed public metrics from 1 member.

Representative image for SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Merged summary

TL;DR - SWE-bench Science is a 119-task benchmark evaluating coding agents on repository-level engineering work across 20 scientific domains. Its strongest tested agent scores below 50% pass@1, exposing persistent difficulties with scientific reasoning, exploration, complete integration, and generalization.

  • The benchmark draws tasks from 98 GitHub repositories and covers issue-driven, expert-exploratory, and engineering-integration scenarios.
  • Claude Code with Opus-5 (max), the best-performing evaluated agent, achieves less than 50% pass@1.
  • Common failures include missing scientific abstractions, superficial or misguided repairs, incomplete system-level coverage, and poor generalization beyond observed cases.
  • Ablations show that well-grounded scientific guidance can improve average performance and token efficiency, while misaligned guidance can anchor agents without improving exact repair success.

Sources (1)

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

arXiv cs.CL Zhipeng Xu, Jiahao Lu, Yining Zheng, Yuxin Wang, Xipeng Qiu 2026-08-20 arXiv:2608.19799
Public signals Hugging Face upvotes 66 · Semantic Scholar citations 2 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 66 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 2 · Influential citations 0 X · N/A Fetched 2026-09-19 14:25:40.137252 UTC

TL;DR - SWE-bench Science is a 119-task benchmark evaluating coding agents on repository-level engineering work across 20 scientific domains. Its strongest tested agent scores below 50% pass@1, exposing persistent difficulties with scientific reasoning, exploration, complete integration, and generalization.

  • The benchmark draws tasks from 98 GitHub repositories and covers issue-driven, expert-exploratory, and engineering-integration scenarios.
  • Claude Code with Opus-5 (max), the best-performing evaluated agent, achieves less than 50% pass@1.
  • Common failures include missing scientific abstractions, superficial or misguided repairs, incomplete system-level coverage, and poor generalization beyond observed cases.
  • Ablations show that well-grounded scientific guidance can improve average performance and token efficiency, while misaligned guidance can anchor agents without improving exact repair success.
item →