🛰️ Daily AI Frontier
‹ back to 2026-08-25

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

Research LLM Agents

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - A controlled 4,400-run pilot tests how finely LLM-agent workflows should be decomposed for cross-border VAT determination. Intermediate decompositions achieved the highest accuracy, but the preregistered evidence threshold was not met, leaving their advantage unconfirmed at pilot scale.

  • Intermediate configurations reached 0.830 accuracy, versus 0.720 and 0.770 for the widest and most fragmented endpoints.
  • A token-matched single agent trailed the leader by 6.5 percentage points, but the confidence interval included zero, so prompt budget may explain the difference.
  • All configurations tolerated availability faults, while a schema-valid hallucinated record degraded every setup and harmed fragmented configurations most.
  • The authors release the oracle, dataset, evaluation harness, raw traces, and analysis pipeline alongside a heuristic that partitions work near the dependency-layer midpoint.

Sources (1)

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

arXiv cs.MA Pedro Santos 2026-08-24 arXiv:2608.23395
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-10 14:23:05.237363 UTC

TL;DR - A controlled 4,400-run pilot tests how finely LLM-agent workflows should be decomposed for cross-border VAT determination. Intermediate decompositions achieved the highest accuracy, but the preregistered evidence threshold was not met, leaving their advantage unconfirmed at pilot scale.

  • Intermediate configurations reached 0.830 accuracy, versus 0.720 and 0.770 for the widest and most fragmented endpoints.
  • A token-matched single agent trailed the leader by 6.5 percentage points, but the confidence interval included zero, so prompt budget may explain the difference.
  • All configurations tolerated availability faults, while a schema-valid hallucinated record degraded every setup and harmed fragmented configurations most.
  • The authors release the oracle, dataset, evaluation harness, raw traces, and analysis pipeline alongside a heuristic that partitions work near the dependency-layer midpoint.
item →