🛰️ Daily AI Frontier
‹ back to 2026-09-19

Score Centering Stabilizes Off-policy Reinforcement Learning

Research LLMs & Foundation Models

Ranking

Overall 83
Content 100
Popularity 43

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper identifies persistent score drift between training and inference engines as a primary cause of instability in off-policy LLM reinforcement learning. Its additive score-centering correction improves stability under engine mismatches without sacrificing rollout efficiency.

  • Score centering cancels bias that otherwise accumulates across training steps.
  • Across models from 0.6B to 30B parameters, it matches or outperforms importance sampling under quantization.
  • Its advantage grows as the training-inference mismatch becomes more severe.
  • The correction composes with importance sampling and improves on pure importance-sampling baselines in staleness experiments.

Sources (1)

Score Centering Stabilizes Off-policy Reinforcement Learning

arXiv cs.LG Martin Marek, Max Ryabinin 2026-09-17 arXiv:2609.20807
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:18:03.556964 UTC

TL;DR - This paper identifies persistent score drift between training and inference engines as a primary cause of instability in off-policy LLM reinforcement learning. Its additive score-centering correction improves stability under engine mismatches without sacrificing rollout efficiency.

  • Score centering cancels bias that otherwise accumulates across training steps.
  • Across models from 0.6B to 30B parameters, it matches or outperforms importance sampling under quantization.
  • Its advantage grows as the training-inference mismatch becomes more severe.
  • The correction composes with importance sampling and improves on pure importance-sampling baselines in staleness experiments.
item →