🛰️ Daily AI Frontier
‹ back to 2026-08-17

Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception

Research LLM Evaluation

Ranking

Overall 75
Content 90
Popularity 41

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper proves that language models’ held-out per-token cross-entropy risk cannot be consistently estimated across all possible data distributions and trained models. It identifies bounded probability flooring and thresholded risk reporting as two qualified remedies.

  • Finite- and infinite-risk states can lie arbitrarily close, preventing any estimator—not only holdout averages—from being universally consistent.
  • The impossibility persists with bounded expected sequence length and full-support models; under full support, inconsistent states are dense.
  • With bounded context, flooring next-token probabilities makes finite risk equivalent to finite expected sequence length.
  • Reporting risk only below a predetermined threshold restores consistency while retaining what model selection requires, but changes the estimation objective.

Sources (1)

Cross-Entropy Risk Estimation for Language Models: Inconsistency Must Be Dense, and the Holdout Method Is No Exception

arXiv cs.LG Hanti Lin 2026-08-16 arXiv:2608.15798
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-30 14:16:44.887090 UTC

TL;DR - This paper proves that language models’ held-out per-token cross-entropy risk cannot be consistently estimated across all possible data distributions and trained models. It identifies bounded probability flooring and thresholded risk reporting as two qualified remedies.

  • Finite- and infinite-risk states can lie arbitrarily close, preventing any estimator—not only holdout averages—from being universally consistent.
  • The impossibility persists with bounded expected sequence length and full-support models; under full support, inconsistent states are dense.
  • With bounded context, flooring next-token probabilities makes finite risk equivalent to finite expected sequence length.
  • Reporting risk only below a predetermined threshold restores consistency while retaining what model selection requires, but changes the estimation objective.
item →