🛰️ Daily AI Frontier
‹ back to 2026-09-07

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

Research LLMs & Foundation Models

Ranking

Overall 85
Content 95
Popularity 62

Observed public metrics from 1 member.

Representative image for RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

Merged summary

TL;DR - RISE is a language-model post-training method that builds a synthetic teacher by extrapolating from the model’s own RLVR training trajectory. It turns sparse outcome rewards into dense token-level supervision without external teachers or privileged conditioning.

  • Extrapolation uses the change between the current checkpoint and a trailing anchor in parameter or output-logit space.
  • RISE alternates RLVR with on-policy distillation: rewards ground reasoning improvements, while the synthetic teacher refines token-level decisions.
  • The teacher is refreshed as the student improves, making distillation a recursive improvement process rather than one-shot compression.
  • Across math, STEM, coding, and multi-turn agentic tasks, RISE reportedly outperforms RLVR-only training and on-policy self-distillation.

Sources (1)

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

arXiv cs.AI Yang Li, Semih Yavuz, Shafiq Joty 2026-09-04 arXiv:2609.05295
Public signals Hugging Face upvotes 17
Providers: Hugging Face · Upvotes 17 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:23:30.293035 UTC

TL;DR - RISE is a language-model post-training method that builds a synthetic teacher by extrapolating from the model’s own RLVR training trajectory. It turns sparse outcome rewards into dense token-level supervision without external teachers or privileged conditioning.

  • Extrapolation uses the change between the current checkpoint and a trailing anchor in parameter or output-logit space.
  • RISE alternates RLVR with on-policy distillation: rewards ground reasoning improvements, while the synthetic teacher refines token-level decisions.
  • The teacher is refreshed as the student improves, making distillation a recursive improvement process rather than one-shot compression.
  • Across math, STEM, coding, and multi-turn agentic tasks, RISE reportedly outperforms RLVR-only training and on-policy self-distillation.
item →