Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
Ranking
Overall
92
Content
100
Popularity
73
Observed public metrics from 1 member.
Merged summary
TL;DR - Organizing synthetic training text into coherent, book-length documents improves language-model mid-training more than isolated rewriting or arbitrary document concatenation. The results identify document-level structure as an important synthetic-data design factor.
- The pipeline produced 686,000 source-grounded textbooks totaling 32B tokens across more than 15,000 disciplines.
- Replacing natural books with this corpus improved downstream performance by 1.09 points on average.
- Content-matched and length-matched controls showed that coherent book packaging—not merely content or document length—drove gains.
- On Llama3-8B, structured books also outperformed randomly concatenated sections and natural books.
Sources (1)
Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
Public signals
Semantic Scholar citations 1 · Semantic Scholar influential citations 0
TL;DR - Organizing synthetic training text into coherent, book-length documents improves language-model mid-training more than isolated rewriting or arbitrary document concatenation. The results identify document-level structure as an important synthetic-data design factor.
- The pipeline produced 686,000 source-grounded textbooks totaling 32B tokens across more than 15,000 disciplines.
- Replacing natural books with this corpus improved downstream performance by 1.09 points on average.
- Content-matched and length-matched controls showed that coherent book packaging—not merely content or document length—drove gains.
- On Llama3-8B, structured books also outperformed randomly concatenated sections and natural books.