🛰️ Daily AI Frontier
‹ back to 2026-07-30

Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

arXiv cs.AI LLMs & Foundation Models Jiawen Tao, Miao Peng, Yaoming Li, Xiaokun Yuan, Mengzhou Wu, Wenhan Yu, Guoan Wang, Nuo Chen, Tong Yang, Maxm Pan 2026-07-30
Representative image for Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

TL;DR - Organizing synthetic training text into coherent, book-length documents improves language-model mid-training more than isolated rewriting or arbitrary document concatenation. The results identify document-level structure as an important synthetic-data design factor.

  • The pipeline produced 686,000 source-grounded textbooks totaling 32B tokens across more than 15,000 disciplines.
  • Replacing natural books with this corpus improved downstream performance by 1.09 points on average.
  • Content-matched and length-matched controls showed that coherent book packaging—not merely content or document length—drove gains.
  • On Llama3-8B, structured books also outperformed randomly concatenated sections and natural books.

view merged work →