🛰️ Daily AI Frontier
‹ back to 2026-07-31

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

Research LLM Agents

Ranking

Overall 79
Content 95
Popularity 43

Observed public metrics from 1 member.

Representative image for ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

Merged summary

TL;DR - ClawTrack evaluates autonomous agents at both outcome and reasoning-trace levels, exposing lucky successes and specific process failures. Its results identify inadequate result verification as a recurring weakness and show that process-based trajectory filtering improves post-training.

  • Includes 320 tasks across 8 domains, 25+ deterministic mock services, and 12,541 task-specific rubric items.
  • Scores each reasoning turn on goal alignment, efficiency, information use, and result verification.
  • Evaluation of 21 models over 16,000+ trials found complementary process dimensions and robustness across judge LLMs.
  • Trace-level scoring enables more precise failure attribution than outcome-only benchmarks.

Sources (1)

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

arXiv cs.LG Xingjian Wu, Xuhang Zhu, Xingchen Liu, Junlin Liu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai 2026-07-30 arXiv:2607.28037
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-30 14:28:54.049608 UTC

TL;DR - ClawTrack evaluates autonomous agents at both outcome and reasoning-trace levels, exposing lucky successes and specific process failures. Its results identify inadequate result verification as a recurring weakness and show that process-based trajectory filtering improves post-training.

  • Includes 320 tasks across 8 domains, 25+ deterministic mock services, and 12,541 task-specific rubric items.
  • Scores each reasoning turn on goal alignment, efficiency, information use, and result verification.
  • Evaluation of 21 models over 16,000+ trials found complementary process dimensions and robustness across judge LLMs.
  • Trace-level scoring enables more precise failure attribution than outcome-only benchmarks.
item →