🛰️ Daily AI Frontier
‹ back to 2026-08-31

A Formal Limitation on Learning Human Language From Textual Corpora

arXiv cs.CL LLMs & Foundation Models Emily Cheng, Ryan Cotterell 2026-08-28

TL;DR - This paper derives information-theoretic limits on recovering a speaker’s intended meaning from text alone, including from LLM hidden states. It shows that ambiguity requiring extralinguistic context cannot be eliminated by better representations, more training data, or additional supervision.

  • Models language as a joint distribution over meanings, contexts, and utterances, then bounds a decoder’s probability of recovering intended meaning.
  • Separates uncertainty into an irreducible component and a component resolvable only through extralinguistic context.
  • Applies to any text featurizer and to both discrete and continuous meaning spaces.
  • Experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference provide empirical support for the theory.

view merged work →