🛰️ Daily AI Frontier
‹ back to 2026-07-31

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

arXiv cs.LG LLM Agents Xingjian Wu, Xuhang Zhu, Xingchen Liu, Junlin Liu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai 2026-07-30
Representative image for ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

TL;DR - ClawTrack evaluates autonomous agents at both outcome and reasoning-trace levels, exposing lucky successes and specific process failures. Its results identify inadequate result verification as a recurring weakness and show that process-based trajectory filtering improves post-training.

  • Includes 320 tasks across 8 domains, 25+ deterministic mock services, and 12,541 task-specific rubric items.
  • Scores each reasoning turn on goal alignment, efficiency, information use, and result verification.
  • Evaluation of 21 models over 16,000+ trials found complementary process dimensions and robustness across judge LLMs.
  • Trace-level scoring enables more precise failure attribution than outcome-only benchmarks.

view merged work →