Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
Ranking
Overall
75
Content
90
Popularity
40
Observed public metrics from 1 member.
Merged summary
TL;DR - This study compares token, byte, and pixel text encodings under controlled linguistic content and downstream capacity. It finds no universally superior encoding; effectiveness depends on the task, languages, capacity, and compute budget.
- Tests verified parallel sentences across 13 languages and five scripts using a shared, variable-width bottleneck.
- Pixels best preserve surface form, while bytes best retain cross-lingual alignment, particularly for same-script languages.
- Tokens perform best for topic classification.
- Sequence length alone does not explain utility: longer representations may preserve highly compressible information, while shorter ones can lose meaning.
Sources (1)
Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This study compares token, byte, and pixel text encodings under controlled linguistic content and downstream capacity. It finds no universally superior encoding; effectiveness depends on the task, languages, capacity, and compute budget.
- Tests verified parallel sentences across 13 languages and five scripts using a shared, variable-width bottleneck.
- Pixels best preserve surface form, while bytes best retain cross-lingual alignment, particularly for same-script languages.
- Tokens perform best for topic classification.
- Sequence length alone does not explain utility: longer representations may preserve highly compressible information, while shorter ones can lose meaning.