Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
- Published
- Source
- arXiv
- Paper number
- 912
- Field
- AI / General
- arXiv ID
- 2608.14290
Key points
- Traditional transformers tie knowledge storage and reasoning compute together layer by layer, so knowledge from other layers has to flow through multiple layers when it is needed.
- Mobius lets every reasoning layer read the same knowledge store and repeatedly refines the internal continuous representation before outputting a token.
- In the 7B comparison, Mobius reaches the same score as a standard transformer trained on the full dataset while using only 62.6% of the data.
- The 35B model achieves an average general-evaluation score of 67.88 and keeps accuracy close to the original Qwen3.5-35B while making question answering about four times faster.
- In one linear-algebra case, it outputs the same answer in 516 tokens, while the original model uses 2,364 tokens, although the reason for the shorter reasoning and higher data efficiency is still only a hypothesis.
Paper links
External research summaries. These are not HDATF publications or measured product results.