🛰️ Daily AI Frontier
‹ back to 2026-07-30

Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

Research LLMs & Foundation Models

Ranking

Overall 92
Content 100
Popularity 73

Observed public metrics from 1 member.

Representative image for Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

Merged summary

TL;DR - Organizing synthetic training text into coherent, book-length documents improves language-model mid-training more than isolated rewriting or arbitrary document concatenation. The results identify document-level structure as an important synthetic-data design factor.

  • The pipeline produced 686,000 source-grounded textbooks totaling 32B tokens across more than 15,000 disciplines.
  • Replacing natural books with this corpus improved downstream performance by 1.09 points on average.
  • Content-matched and length-matched controls showed that coherent book packaging—not merely content or document length—drove gains.
  • On Llama3-8B, structured books also outperformed randomly concatenated sections and natural books.

Sources (1)

Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training

arXiv cs.AI Jiawen Tao, Miao Peng, Yaoming Li, Xiaokun Yuan, Mengzhou Wu, Wenhan Yu, Guoan Wang, Nuo Chen, Tong Yang, Maxm Pan 2026-07-30 arXiv:2607.28109
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-08-24 14:30:19.981551 UTC

TL;DR - Organizing synthetic training text into coherent, book-length documents improves language-model mid-training more than isolated rewriting or arbitrary document concatenation. The results identify document-level structure as an important synthetic-data design factor.

  • The pipeline produced 686,000 source-grounded textbooks totaling 32B tokens across more than 15,000 disciplines.
  • Replacing natural books with this corpus improved downstream performance by 1.09 points on average.
  • Content-matched and length-matched controls showed that coherent book packaging—not merely content or document length—drove gains.
  • On Llama3-8B, structured books also outperformed randomly concatenated sections and natural books.
item →