BAT 集体重做「AI 预训练」,补「脏数据」的坑
TL;DR - Chinese foundation-model developers, including major technology companies, are reportedly redoing pretraining pipelines after low-quality, duplicated, mislabeled, or evaluation-contaminated data impaired model performance. The shift underscores that data governance and feedback pipelines—not just model scale and compute—are becoming critical competitive infrastructure.
- Early quantity-focused collection introduced web spam, synthetic duplicates, unclear labeling, and benchmark leakage, wasting compute and sometimes forcing complete retraining.
- Teams are shifting toward information density, cleanliness, iterative validation, and closer coordination between data and pretraining groups rather than volume-based data KPIs.
- High-quality expert examples and long-horizon agent trajectories remain expensive and capacity-constrained, prompting some firms to seek interaction data through API relays and model-routing platforms.
- Most companies have not established a true data flywheel because organizational silos, scarce data-engineering talent, and limited domestic coding and office-agent usage impede the flow of user interactions back into training.