Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data
TL;DR - This paper uses Fisher-Rao information geometry to derive theoretical guarantees for the minimum proportion of fresh human data needed to prevent model collapse during recursive training on synthetic data. Its bounds remain meaningful in high-dimensional categorical distributions, unlike prior Euclidean analyses.
- Models recursively trained on synthetic data can progressively lose fidelity to the underlying real-data distribution.
- The analysis models training dynamics on the probability simplex using the Fisher-Rao metric.
- It derives quantitative contraction and invariance bounds that do not become trivial as dimensionality increases.
- The results indicate that the human-to-synthetic data ratio required for stable training differs from previous estimates.