Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - KOPA-Bench evaluates multi-step LLM tool use across live Korean government APIs, while EDGE synthesizes executable training trajectories from API connections verified through real calls. Using EDGE data and GRPO, a fine-tuned 9B model nearly matches an untuned 27B model from the same family.
- KOPA-Bench contains 145 real-world tasks targeting multi-step tool calling in open-source, on-premise agents.
- EDGE constructs a graph linking tool outputs to compatible inputs and retains only connections that execute successfully against live APIs.
- Traversing these verified links produces grounded, executable multi-step training trajectories.
- Improvements transfer beyond KOPA-Bench to the BFCL tool-calling benchmark.
Sources (1)
Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - KOPA-Bench evaluates multi-step LLM tool use across live Korean government APIs, while EDGE synthesizes executable training trajectories from API connections verified through real calls. Using EDGE data and GRPO, a fine-tuned 9B model nearly matches an untuned 27B model from the same family.
- KOPA-Bench contains 145 real-world tasks targeting multi-step tool calling in open-source, on-premise agents.
- EDGE constructs a graph linking tool outputs to compatible inputs and retains only connections that execute successfully against live APIs.
- Traversing these verified links produces grounded, executable multi-step training trajectories.
- Improvements transfer beyond KOPA-Bench to the BFCL tool-calling benchmark.