🛰️ Daily AI Frontier
‹ back to 2026-09-07

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

Research LLM Agents

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Representative image for Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

Merged summary

TL;DR - KOPA-Bench evaluates multi-step LLM tool use across live Korean government APIs, while EDGE synthesizes executable training trajectories from API connections verified through real calls. Using EDGE data and GRPO, a fine-tuned 9B model nearly matches an untuned 27B model from the same family.

  • KOPA-Bench contains 145 real-world tasks targeting multi-step tool calling in open-source, on-premise agents.
  • EDGE constructs a graph linking tool outputs to compatible inputs and retains only connections that execute successfully against live APIs.
  • Traversing these verified links produces grounded, executable multi-step training trajectories.
  • Improvements transfer beyond KOPA-Bench to the BFCL tool-calling benchmark.

Sources (1)

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

arXiv cs.AI Dain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh, Kyuseong Lim 2026-09-04 arXiv:2609.05395
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-09 08:03:55.133518 UTC

TL;DR - KOPA-Bench evaluates multi-step LLM tool use across live Korean government APIs, while EDGE synthesizes executable training trajectories from API connections verified through real calls. Using EDGE data and GRPO, a fine-tuned 9B model nearly matches an untuned 27B model from the same family.

  • KOPA-Bench contains 145 real-world tasks targeting multi-step tool calling in open-source, on-premise agents.
  • EDGE constructs a graph linking tool outputs to compatible inputs and retains only connections that execute successfully against live APIs.
  • Traversing these verified links produces grounded, executable multi-step training trajectories.
  • Improvements transfer beyond KOPA-Bench to the BFCL tool-calling benchmark.
item →