1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
Ranking
Overall
85
Content
95
Popularity
61
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper introduces an information-efficiency ratio (IER) for selecting which tokens receive teacher supervision during on-policy distillation. By considering both token usefulness and gradient-estimation reliability, it matches or surpasses full supervision on reasoning tasks while supervising only 0.1%–1% of tokens.
- IER measures relative gradient-estimation error using a signal-to-noise decomposition and an optimal scalar baseline.
- A candidate-set approximation makes IER practical for token selection while preserving the sampled reverse-KL objective.
- Adding IER improves existing token selectors across multiple mathematical and medical reasoning settings.
- The results suggest sparse distillation should prioritize reliable gradients as well as useful teacher guidance.
Sources (1)
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
Public signals
Hugging Face upvotes 16
TL;DR - This paper introduces an information-efficiency ratio (IER) for selecting which tokens receive teacher supervision during on-policy distillation. By considering both token usefulness and gradient-estimation reliability, it matches or surpasses full supervision on reasoning tasks while supervising only 0.1%–1% of tokens.
- IER measures relative gradient-estimation error using a signal-to-noise decomposition and an optimal scalar baseline.
- A candidate-set approximation makes IER practical for token selection while preserving the sampled reverse-KL objective.
- Adding IER improves existing token selectors across multiple mathematical and medical reasoning settings.
- The results suggest sparse distillation should prioritize reliable gradients as well as useful teacher guidance.