Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

Published
Source
arXiv
Paper number
912
Field
AI / General
arXiv ID
2608.14290

Key points

  • Traditional transformers tie knowledge storage and reasoning compute together layer by layer, so knowledge from other layers has to flow through multiple layers when it is needed.
  • Mobius lets every reasoning layer read the same knowledge store and repeatedly refines the internal continuous representation before outputting a token.
  • In the 7B comparison, Mobius reaches the same score as a standard transformer trained on the full dataset while using only 62.6% of the data.
  • The 35B model achieves an average general-evaluation score of 67.88 and keeps accuracy close to the original Qwen3.5-35B while making question answering about four times faster.
  • In one linear-algebra case, it outputs the same answer in 516 tokens, while the original model uses 2,364 tokens, although the reason for the shorter reasoning and higher data efficiency is still only a hypothesis.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)