🛰️ Daily AI Frontier
‹ back to 2026-08-07

中国版「生物DeepSeek」诞生!4个牛津学霸,让AI接管生命科学

WeChat: 新智元 Bioinformatics AI 2026-08-04
Representative image for 中国版「生物DeepSeek」诞生!4个牛津学霸,让AI接管生命科学

TL;DR - Shenzhen-based startup 津渡生科 (Jindu/BioFord), founded by four Oxford-trained returnees, has released GeneLLM, a multi-omics foundation model pretrained directly on raw sequencing data, plus a robotic lab-automation stack — pitched as China's "biology DeepSeek" after publications in Nature Communications and Advanced Science.

  • GeneLLM tokenizes ~150bp RNA-seq reads via 7-mer sliding windows and does next-base prediction with a Transformer, skipping gene annotation, alignment, and human labels; training is two-stage (unsupervised pretraining + prototype mining, then patient-level "Disease Tuning").
  • Scale claimed: 1.5B params over 3.5T bases, plus a 30B-param XLarge version, trained on ~tens of trillions of RNA reads on a ~100-GPU NVIDIA A100 cluster.
  • Efficiency claim: maintains AUC > 0.8 at 1Gb ultra-shallow sequencing depth vs. the conventional 6Gb, cited as an ~83% cost reduction for disease detection.
  • Beyond the model: "BioFord Harness" compiles experiment DSLs into instrument commands via a Universal Instrument Abstraction Layer (PCR, plate readers, flow cytometers, liquid handlers), with five agents (literature, experiment design, science, scheduling, data analysis) closing a DBTL loop; company reports four funding rounds in one year (Sequoia China Seed angel+ through a ~100M RMB Series A led by 高特佳).
  • Note: all performance and deployment figures come from the company/media article, not independently verified here.

view merged work →