Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

Published
Source
arXiv
Paper number
590
Field
AI / General
arXiv ID
2607.08393

Key points

  • It formalizes the Knowing-Using Gap, the phenomenon in which memorization is nearly perfect but generalization, especially multi-step reasoning, lags far behind.
  • It introduces self-patching, a causal diagnostic tool that recovers generalization failures immediately by moving a representation from one layer to another.
  • It proposes the knowledge-circuit misalignment hypothesis, where memorized representations exist in early and late layers but are not routed into the middle reasoning layers, causing failure.
  • The paper validates this across six models, from Qwen-2.5 1.5B to 7B and LLaMA-3.2 1B to 8B, and across two domains, biomedical and academic.
  • A fixed heuristic using only two layer-pair settings, 0.8L to 0.5L and 0.1L to 0.5L, recovers 58 to 75 percent of the oracle improvement.
  • Self-patching is far better than CoT prompting or irrelevant patching, which supports the claim that the root cause is structural.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)