Language models suffer from a curse of ambiguity
Ranking
Overall
82
Content
100
Popularity
41
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper identifies a “curse of ambiguity”: neural networks learn uncertain next-token probability distributions less accurately than concentrated ones. This matters as LLM training increasingly depends on sampling from learned distributions.
- Ambiguous distributions require greater model capacity and larger embeddings to store and represent accurately.
- They take more optimization steps to fit and amplify noise during token sampling.
- Controlled synthetic experiments support the theory, while models trained on real data show similar patterns.
- The analysis offers a framework for judging when an LLM’s output distribution is trustworthy.
Sources (1)
Language models suffer from a curse of ambiguity
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper identifies a “curse of ambiguity”: neural networks learn uncertain next-token probability distributions less accurately than concentrated ones. This matters as LLM training increasingly depends on sampling from learned distributions.
- Ambiguous distributions require greater model capacity and larger embeddings to store and represent accurately.
- They take more optimization steps to fit and amplify noise during token sampling.
- Controlled synthetic experiments support the theory, while models trained on real data show similar patterns.
- The analysis offers a framework for judging when an LLM’s output distribution is trustworthy.