🛰️ Daily AI Frontier
‹ back to 2026-08-17

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

Research LLMs & Foundation Models

Ranking

Overall 87
Content 95
Popularity 68

Observed public metrics from 1 member.

Representative image for Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

Merged summary

TL;DR - Mobius-v0 separates knowledge storage in a shared FFN memory from iterative reasoning in multiple self-attention modules. The approach reportedly matches Transformer baselines with less training data or nearly 4Ă— faster inference.

  • Reasoners repeatedly query shared memory using hidden states as both cache and information carrier.
  • A from-scratch 7B model matches a 7B Transformer baseline’s downstream score using 62.6% of its training data.
  • Intern-S2-Mobius, continually pretrained from Qwen3.5-35B, achieves a similar downstream score with nearly 4Ă— end-to-end inference speedup.

Sources (1)

Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

arXiv cs.AI Kai Chen, Jifeng Ding, Ning Ding, Jiaye Ge, Lixin Gu, Yicheng Gu, Qipeng Guo, Ermo Hua, Haian Huang, Haozheng Hou, Jie Hou, Xiangyu Hong, Che Jiang, Minxi Jin, Cheng Liang, Dahua Lin, Dawei Liu, Kuikun Liu, Chengqi Lv, Haijun Lv, Han Lv, Ningsheng Ma, Biqing Qi, Jianmin Qian, Shiya Su, Youbang Sun, Huanze Tang, Zhongbo Tian, Hanjing Wang, Rui Wang, Ting Wang, Yi Wang, Baiting Wu, Jun Xu, Bowen Yang, Hui Wang, Weida Wang, Haochen Ye, Jiashuo Yu, Shan Yu, Xiaoyi Yu, Qirui Zeng, Qi Zhang, Ming Zhang, Wenwei Zhang, Bowen Zhou, Xinyu Zhou 2026-08-14 arXiv:2608.14290
Public signals Hugging Face upvotes 34
Providers: Hugging Face · Upvotes 34 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-15 14:32:42.318468 UTC

TL;DR - Mobius-v0 separates knowledge storage in a shared FFN memory from iterative reasoning in multiple self-attention modules. The approach reportedly matches Transformer baselines with less training data or nearly 4Ă— faster inference.

  • Reasoners repeatedly query shared memory using hidden states as both cache and information carrier.
  • A from-scratch 7B model matches a 7B Transformer baseline’s downstream score using 62.6% of its training data.
  • Intern-S2-Mobius, continually pretrained from Qwen3.5-35B, achieves a similar downstream score with nearly 4Ă— end-to-end inference speedup.
item →