Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks
Merged summary
TL;DR - This paper introduces an LLM-based framework and company-isolated benchmark for detecting financial statement fraud from structured financial data and summarized MD&A text. The evaluation targets generalization to unseen companies, avoiding overly optimistic random splits.
- Proposes Company-Isolated FSFD (CI-FSFD) as a more realistic benchmark task.
- Publishes a U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels.
- Integrates structured financial metrics with unstructured report text using LLMs.
- Reports state-of-the-art performance on CI-FSFD and highlights the value of textual data.
Sources (1)
Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks
TL;DR - This paper introduces an LLM-based framework and company-isolated benchmark for detecting financial statement fraud from structured financial data and summarized MD&A text. The evaluation targets generalization to unseen companies, avoiding overly optimistic random splits.
- Proposes Company-Isolated FSFD (CI-FSFD) as a more realistic benchmark task.
- Publishes a U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels.
- Integrates structured financial metrics with unstructured report text using LLMs.
- Reports state-of-the-art performance on CI-FSFD and highlights the value of textual data.