🛰️ Daily AI Frontier
‹ back to 2026-09-24

BAT 集体重做「AI 预训练」,补「脏数据」的坑

雷峰网 (AI科技评论) LLMs & Foundation Models 2026-09-24
Representative image for BAT 集体重做「AI 预训练」,补「脏数据」的坑

TL;DR - Chinese foundation-model developers, including major technology companies, are reportedly redoing pretraining pipelines after low-quality, duplicated, mislabeled, or evaluation-contaminated data impaired model performance. The shift underscores that data governance and feedback pipelines—not just model scale and compute—are becoming critical competitive infrastructure.

  • Early quantity-focused collection introduced web spam, synthetic duplicates, unclear labeling, and benchmark leakage, wasting compute and sometimes forcing complete retraining.
  • Teams are shifting toward information density, cleanliness, iterative validation, and closer coordination between data and pretraining groups rather than volume-based data KPIs.
  • High-quality expert examples and long-horizon agent trajectories remain expensive and capacity-constrained, prompting some firms to seek interaction data through API relays and model-routing platforms.
  • Most companies have not established a true data flywheel because organizational silos, scarce data-engineering talent, and limited domestic coding and office-agent usage impede the flow of user interactions back into training.

view merged work →