🛰️ Daily AI Frontier
‹ back to 2026-09-02

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

Research LLM Agents

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - WorldBench is a multilingual benchmark of 1,600 culturally and persona-grounded workflows for evaluating agents in sandboxed environments. Frontier models achieve only 49.2% Constrained Task Success, exposing brittleness in long-horizon tasks and preserving environment state.

  • Covers seven languages and eight cultures, with tasks refined by language- and culture-specific human annotators.
  • Evaluates realistic multi-step workflows through structured actions in sandboxed environments.
  • Introduces Constrained Task Success (CTS), combining task completion, minimal modification, and complementary metrics via deterministic and LLM-based judging.
  • All evaluated models show substantial gaps between task correctness and environment preservation.

Sources (1)

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

arXiv cs.AI Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch 2026-09-01 arXiv:2609.01056
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-21 14:26:38.491349 UTC

TL;DR - WorldBench is a multilingual benchmark of 1,600 culturally and persona-grounded workflows for evaluating agents in sandboxed environments. Frontier models achieve only 49.2% Constrained Task Success, exposing brittleness in long-horizon tasks and preserving environment state.

  • Covers seven languages and eight cultures, with tasks refined by language- and culture-specific human annotators.
  • Evaluates realistic multi-step workflows through structured actions in sandboxed environments.
  • Introduces Constrained Task Success (CTS), combining task completion, minimal modification, and complementary metrics via deterministic and LLM-based judging.
  • All evaluated models show substantial gaps between task correctness and environment preservation.
item →