WorldBench: Culturally Grounded Benchmark for Multilingual Agents
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - WorldBench is a multilingual benchmark of 1,600 culturally and persona-grounded workflows for evaluating agents in sandboxed environments. Frontier models achieve only 49.2% Constrained Task Success, exposing brittleness in long-horizon tasks and preserving environment state.
- Covers seven languages and eight cultures, with tasks refined by language- and culture-specific human annotators.
- Evaluates realistic multi-step workflows through structured actions in sandboxed environments.
- Introduces Constrained Task Success (CTS), combining task completion, minimal modification, and complementary metrics via deterministic and LLM-based judging.
- All evaluated models show substantial gaps between task correctness and environment preservation.
Sources (1)
WorldBench: Culturally Grounded Benchmark for Multilingual Agents
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - WorldBench is a multilingual benchmark of 1,600 culturally and persona-grounded workflows for evaluating agents in sandboxed environments. Frontier models achieve only 49.2% Constrained Task Success, exposing brittleness in long-horizon tasks and preserving environment state.
- Covers seven languages and eight cultures, with tasks refined by language- and culture-specific human annotators.
- Evaluates realistic multi-step workflows through structured actions in sandboxed environments.
- Introduces Constrained Task Success (CTS), combining task completion, minimal modification, and complementary metrics via deterministic and LLM-based judging.
- All evaluated models show substantial gaps between task correctness and environment preservation.