🛰️ Daily AI Frontier
‹ back to 2026-09-25

Rufus-Air: An Open LLM Post-Training Recipe

Research LLMs & Foundation Models

Ranking

Overall 89
Content 100
Popularity 64

Observed public metrics from 1 member.

Merged summary

TL;DR - Rufus-Air presents an open, reproducible eight-stage post-training recipe for the 106B-parameter GLM-4.5-Air-Base model. It demonstrates that carefully ordered training stages, reliable rewards, prompt filtering, and infrastructure choices can produce results competitive with similarly sized open models.

  • The pipeline combines SFT, specialized reasoning/coding/instruction-following RL, agent training, and final RLHF.
  • Stages move from foundational to advanced capabilities and from verifiable rewards toward softer judge-based signals.
  • Diverse high-quality SFT establishes the capability baseline, while difficulty filtering keeps RL prompts within a productive range.
  • The recipe uses open-source components and public data without new human annotation or an internal distillation teacher.

Sources (1)

Rufus-Air: An Open LLM Post-Training Recipe

arXiv cs.CL Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He, Hongye Jin, Linwei Li, Shiyang Li, Fenglin Liu, Xin Liu, Priyanka Nigam, Haoyang Wen, Zhenghao Xu, Zhuocheng Xu, Bing Yin, Qingyu Yin, Chao Zhang, Rongzhi Zhang, Zhihan Zhang, Zixuan Zhang, Zixuan Zhang, Tuo Zhao 2026-09-24 arXiv:2609.29421
Public signals Hugging Face upvotes 11
Providers: Hugging Face · Upvotes 11 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:11.454674 UTC

TL;DR - Rufus-Air presents an open, reproducible eight-stage post-training recipe for the 106B-parameter GLM-4.5-Air-Base model. It demonstrates that carefully ordered training stages, reliable rewards, prompt filtering, and infrastructure choices can produce results competitive with similarly sized open models.

  • The pipeline combines SFT, specialized reasoning/coding/instruction-following RL, agent training, and final RLHF.
  • Stages move from foundational to advanced capabilities and from verifiable rewards toward softer judge-based signals.
  • Diverse high-quality SFT establishes the capability baseline, while difficulty filtering keeps RL prompts within a productive range.
  • The recipe uses open-source components and public data without new human annotation or an internal distillation teacher.
item →