🛰️ Daily AI Frontier
‹ back to 2026-08-24

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

Research LLM Agents

Ranking

Overall 88
Content 100
Popularity 59

Observed public metrics from 1 member.

Representative image for EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

Merged summary

TL;DR - EarthVerse is a benchmark for evaluating tool-using scientific agents on 405 reproducible Earth-system and natural-hazard investigations. Results show that even strong systems struggle to maintain reliable end-to-end reasoning across heterogeneous evidence, calculations, units, and physical interpretation.

  • Tasks span 199 documented events and 19 hazard families, with provenance-preserving answers and executable ground truth.
  • The benchmark evaluates evidence selection, tool use, memory, reasoning, interaction, and scientific execution under a controlled protocol.
  • Across 25 model and agent systems, the best mean answer-unit accuracy was 84.65%, but the highest Strict@95 score was only 34.81%.
  • This gap indicates that agents often solve individual steps correctly while failing to preserve a consistent scientific chain across the full investigation.

Sources (1)

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

arXiv cs.AI Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong, Weidong Tang, Yuxuan Liang, Weijia Li, Ming Jin, Shirui Pan, Yuhao Kang, Dingyi Zhuang, Jinhua Zhao 2026-08-24 arXiv:2608.23525
Public signals Hugging Face upvotes 1
Providers: Hugging Face · Upvotes 1 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-22 14:32:27.594758 UTC

TL;DR - EarthVerse is a benchmark for evaluating tool-using scientific agents on 405 reproducible Earth-system and natural-hazard investigations. Results show that even strong systems struggle to maintain reliable end-to-end reasoning across heterogeneous evidence, calculations, units, and physical interpretation.

  • Tasks span 199 documented events and 19 hazard families, with provenance-preserving answers and executable ground truth.
  • The benchmark evaluates evidence selection, tool use, memory, reasoning, interaction, and scientific execution under a controlled protocol.
  • Across 25 model and agent systems, the best mean answer-unit accuracy was 84.65%, but the highest Strict@95 score was only 34.81%.
  • This gap indicates that agents often solve individual steps correctly while failing to preserve a consistent scientific chain across the full investigation.
item →