Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
TL;DR - Organizing synthetic training text into coherent, book-length documents improves language-model mid-training more than isolated rewriting or arbitrary document concatenation. The results identify document-level structure as an important synthetic-data design factor.
- The pipeline produced 686,000 source-grounded textbooks totaling 32B tokens across more than 15,000 disciplines.
- Replacing natural books with this corpus improved downstream performance by 1.09 points on average.
- Content-matched and length-matched controls showed that coherent book packaging—not merely content or document length—drove gains.
- On Llama3-8B, structured books also outperformed randomly concatenated sections and natural books.