🛰️ Daily AI Frontier
‹ back to 2026-08-04

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

arXiv cs.CL LLM Agents Abdulrahman AlRabah, Xiaocheng Yang, Dilek Hakkani-TĂĽr, Abdussalam Alawini 2026-08-03
Representative image for PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

TL;DR - PredAct-Bench is a benchmark for tool-augmented dialogue agents that must reason with statistically imperfect prediction tools while advising a human decision-maker, using education as a measurable testbed. It matters because real deployments in healthcare, finance, and education involve noisy tools, yet existing benchmarks assume perfect tool outputs.

  • Builds an AI-assisted human decision-making benchmark where the agent uses noisy predictors to guide a user, grounded in domains with available ground-truth outcomes and clear intervention decisions.
  • Introduces episode-level Relative AI-Reliance (RAIR) and Relative Self-Reliance (RSR) metrics, extending prior trust-calibration frameworks from single decisions to multi-turn dialogue.
  • Evaluates 13 closed- and open-source SOTA LLMs on two educational datasets — OULAD (real UK Open University assessment trajectories) and PREDACT-CS (60 courses with real final grades plus synthetic weekly score trajectories) — plus a human study with instructors and TAs.
  • Key finding: under noisy tools, current models fail to surface adequate visibility/uncertainty to teachers, leaving users prone to over-reliance on wrong suggestions and hallucinations.

view merged work →