ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
TL;DR - ClawTrack evaluates autonomous agents at both outcome and reasoning-trace levels, exposing lucky successes and specific process failures. Its results identify inadequate result verification as a recurring weakness and show that process-based trajectory filtering improves post-training.
- Includes 320 tasks across 8 domains, 25+ deterministic mock services, and 12,541 task-specific rubric items.
- Scores each reasoning turn on goal alignment, efficiency, information use, and result verification.
- Evaluation of 21 models over 16,000+ trials found complementary process dimensions and robustness across judge LLMs.
- Trace-level scoring enables more precise failure attribution than outcome-only benchmarks.