🛰️ Daily AI Frontier
‹ back to 2026-08-31

A Formal Limitation on Learning Human Language From Textual Corpora

Research LLMs & Foundation Models

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper derives information-theoretic limits on recovering a speaker’s intended meaning from text alone, including from LLM hidden states. It shows that ambiguity requiring extralinguistic context cannot be eliminated by better representations, more training data, or additional supervision.

  • Models language as a joint distribution over meanings, contexts, and utterances, then bounds a decoder’s probability of recovering intended meaning.
  • Separates uncertainty into an irreducible component and a component resolvable only through extralinguistic context.
  • Applies to any text featurizer and to both discrete and continuous meaning spaces.
  • Experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference provide empirical support for the theory.

Sources (1)

A Formal Limitation on Learning Human Language From Textual Corpora

arXiv cs.CL Emily Cheng, Ryan Cotterell 2026-08-28 arXiv:2608.28560
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-16 14:17:26.346002 UTC

TL;DR - This paper derives information-theoretic limits on recovering a speaker’s intended meaning from text alone, including from LLM hidden states. It shows that ambiguity requiring extralinguistic context cannot be eliminated by better representations, more training data, or additional supervision.

  • Models language as a joint distribution over meanings, contexts, and utterances, then bounds a decoder’s probability of recovering intended meaning.
  • Separates uncertainty into an irreducible component and a component resolvable only through extralinguistic context.
  • Applies to any text featurizer and to both discrete and continuous meaning spaces.
  • Experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference provide empirical support for the theory.
item →