🛰️ Daily AI Frontier
‹ back to 2026-07-16

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

Research LLMs & Foundation Models

Ranking

Overall 83
Content 90
Popularity 67

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper tests whether LLM in-context estimates obey the law of total probability, finding widespread violations of statistical self-consistency—a reference-free way to evaluate frontier models.

  • Uses binary trees to recursively partition populations into finer subpopulations, prompts LLMs with verbalized subpopulation descriptions, then aggregates estimates back to population level and compares across partition granularities.
  • Across multiple domains and state-of-the-art models, LLM estimates broadly violate basic probabilistic identities (prior-weighted conditionals should aggregate into marginals).
  • Identifies a "macro fallacy": estimates reconstructed from fine-grained subpopulation responses often align better with human reference data than direct population-level estimates—models hold subpopulation knowledge but fail to propagate it into aggregates.
  • The effect persists across tree structures and tasks and is partially recoverable via implicit prompting, positioning self-consistency as an unsaturated, reference-free evaluation criterion.

Sources (1)

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

arXiv cs.CL Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-DĂĽnner 2026-07-16 arXiv:2607.15277
Public signals Hugging Face upvotes 10
Providers: Hugging Face · Upvotes 10 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-15 14:33:55.115865 UTC

TL;DR - This paper tests whether LLM in-context estimates obey the law of total probability, finding widespread violations of statistical self-consistency—a reference-free way to evaluate frontier models.

  • Uses binary trees to recursively partition populations into finer subpopulations, prompts LLMs with verbalized subpopulation descriptions, then aggregates estimates back to population level and compares across partition granularities.
  • Across multiple domains and state-of-the-art models, LLM estimates broadly violate basic probabilistic identities (prior-weighted conditionals should aggregate into marginals).
  • Identifies a "macro fallacy": estimates reconstructed from fine-grained subpopulation responses often align better with human reference data than direct population-level estimates—models hold subpopulation knowledge but fail to propagate it into aggregates.
  • The effect persists across tree structures and tasks and is partially recoverable via implicit prompting, positioning self-consistency as an unsaturated, reference-free evaluation criterion.
item →