「用初中数学讲明白AI」第1章:一道填空题值一万亿美元
TL;DR - An introductory explainer uses basic probability to describe LLMs as next-token predictors. It argues that their capabilities—and limitations—emerge largely from scaling this simple objective across vast models and datasets.
- LLMs produce text iteratively by predicting a probability distribution over possible next tokens.
- Scale distinguishes modern LLMs from basic autocomplete: far more parameters, training data, vocabulary, and usable context.
- The article situates LLMs within AI ⊃ machine learning ⊃ deep learning ⊃ large models.
- Next-token prediction can generate fluent text but does not guarantee factual accuracy, mathematical reasoning, memory, or tool access.