More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It
Ranking
Overall
82
Content
100
Popularity
41
Observed public metrics from 1 member.
Merged summary
TL;DR - Power Sampling can increase probability mass on correct reasoning trajectories yet reduce downstream accuracy by narrowing useful path coverage. A deformation-controlled, support-preserving alternative avoids this failure and improves multi-sample reasoning inference.
- Standard Power Sampling caused self-consistency accuracy drops of up to 18.5 percentage points.
- Fixed exponents create a “dose mismatch,” changing distributions unevenly across problems.
- Global sharpening creates a “coverage mismatch,” suppressing moderate-probability reasoning paths despite high pass@k.
- Weighted self-consistency with the repaired sampler reversed these losses under the same inference budget.
Sources (1)
More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Power Sampling can increase probability mass on correct reasoning trajectories yet reduce downstream accuracy by narrowing useful path coverage. A deformation-controlled, support-preserving alternative avoids this failure and improves multi-sample reasoning inference.
- Standard Power Sampling caused self-consistency accuracy drops of up to 18.5 percentage points.
- Fixed exponents create a “dose mismatch,” changing distributions unevenly across problems.
- Global sharpening creates a “coverage mismatch,” suppressing moderate-probability reasoning paths despite high pass@k.
- Weighted self-consistency with the repaired sampler reversed these losses under the same inference budget.