🛰️ Daily AI Frontier
‹ back to 2026-08-31

NumBench: Diagnosing Counting Failures in Text-to-Image Models

Research Multimodal & Generative

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Representative image for NumBench: Diagnosing Counting Failures in Text-to-Image Models

Merged summary

TL;DR - NumBench is a 640,000-prompt benchmark designed to diagnose object-counting failures in text-to-image models. Tests across nine methods show performance falls sharply as requested counts rise, with every evaluated method struggling above 50 objects.

  • NumBench covers 1,600 object categories and counts from 1 to 100, systematically varying composition, spatial guidance, and appearance.
  • The proposed process model treats object instances as competing for finite resolvable image regions, predicting counting deficits from collisions and improvements from coordinated placement.
  • Count range had the largest measured effect on accuracy, followed by layout and composition; grid guidance performed best among guided layouts.
  • A confidence-weighted metric combining three calibrated detectors was supported by a 14,400-image human study through count 50, with additional results indicating transfer to natural-language prompts.

Sources (1)

NumBench: Diagnosing Counting Failures in Text-to-Image Models

arXiv cs.CV Sandeep Wadhwa, Mayank Vatsa, Richa Singh, Parrva Chirag Shah, Prakhar Galriya 2026-08-28 arXiv:2608.28206
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-03 14:15:43.418970 UTC

TL;DR - NumBench is a 640,000-prompt benchmark designed to diagnose object-counting failures in text-to-image models. Tests across nine methods show performance falls sharply as requested counts rise, with every evaluated method struggling above 50 objects.

  • NumBench covers 1,600 object categories and counts from 1 to 100, systematically varying composition, spatial guidance, and appearance.
  • The proposed process model treats object instances as competing for finite resolvable image regions, predicting counting deficits from collisions and improvements from coordinated placement.
  • Count range had the largest measured effect on accuracy, followed by layout and composition; grid guidance performed best among guided layouts.
  • A confidence-weighted metric combining three calibrated detectors was supported by a 14,400-image human study through count 50, with additional results indicating transfer to natural-language prompts.
item →