Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
- Published
- Source
- arXiv
- Paper number
- 038
- Field
- Architecture / Reasoning
- arXiv ID
- 2502.13842
Key points
- Small large language models, or LLMs, face an inherent performance bottleneck when handling complex reasoning tasks because of fixed compute budgets and limited parameter counts.
- The traditional approach of continually scaling LLM parameters to improve performance runs into diminishing returns and rapidly rising compute and deployment costs, raising sustainability concerns.
- Empirical analysis shows that "hard" tokens induce persistent gradient oscillations and sharp spikes across Transformer layers, indicating structural burden and insufficient depth for processing complex information.
- It introduces the concept of an Inner Thinking Step, which breaks the generation of a single token into a series of internal reasoning steps and repeatedly refines its representation.
- It implements Residual Thinking Connections, or RTC, which accumulate the output of each thinking step to enable gradual representation refinement and break through performance bottlenecks.
- It applies Adaptive Token Routing, or ATR, to dynamically assign extra thinking steps only to the most important tokens, avoiding unnecessary work on easy tokens and improving compute efficiency, while Thinking Step Encoding, or TSE, provides contextual information at each step.
Paper links
External research summaries. These are not HDATF publications or measured product results.