NumBench: Diagnosing Counting Failures in Text-to-Image Models
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - NumBench is a 640,000-prompt benchmark designed to diagnose object-counting failures in text-to-image models. Tests across nine methods show performance falls sharply as requested counts rise, with every evaluated method struggling above 50 objects.
- NumBench covers 1,600 object categories and counts from 1 to 100, systematically varying composition, spatial guidance, and appearance.
- The proposed process model treats object instances as competing for finite resolvable image regions, predicting counting deficits from collisions and improvements from coordinated placement.
- Count range had the largest measured effect on accuracy, followed by layout and composition; grid guidance performed best among guided layouts.
- A confidence-weighted metric combining three calibrated detectors was supported by a 14,400-image human study through count 50, with additional results indicating transfer to natural-language prompts.
Sources (1)
NumBench: Diagnosing Counting Failures in Text-to-Image Models
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - NumBench is a 640,000-prompt benchmark designed to diagnose object-counting failures in text-to-image models. Tests across nine methods show performance falls sharply as requested counts rise, with every evaluated method struggling above 50 objects.
- NumBench covers 1,600 object categories and counts from 1 to 100, systematically varying composition, spatial guidance, and appearance.
- The proposed process model treats object instances as competing for finite resolvable image regions, predicting counting deficits from collisions and improvements from coordinated placement.
- Count range had the largest measured effect on accuracy, followed by layout and composition; grid guidance performed best among guided layouts.
- A confidence-weighted metric combining three calibrated detectors was supported by a 14,400-image human study through count 50, with additional results indicating transfer to natural-language prompts.