ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance
TL;DR - ReguSim and ReguBench evaluate whether financial-market LLM agents ground their actions and monitoring judgments in executable compliance rules and reliable evidence. The results show that stated rationales are insufficient for auditing compliance and that structured baselines can outperform prompt-only LLM monitors.
- ReguSim separates stated reasoning, attempted actions, execution enforcement, and monitoring evidence.
- Visible rules reduced but did not eliminate rejected trades; incentive and persona framing also changed agent behavior.
- Trader rationales could mislead independent monitors unless execution-enforcement evidence was provided.
- Simple structured monitoring baselines matched or exceeded prompt-only LLMs.