Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
- Published
- Source
- arXiv
- Paper number
- 590
- Field
- AI / General
- arXiv ID
- 2607.08393
Key points
- It formalizes the Knowing-Using Gap, the phenomenon in which memorization is nearly perfect but generalization, especially multi-step reasoning, lags far behind.
- It introduces self-patching, a causal diagnostic tool that recovers generalization failures immediately by moving a representation from one layer to another.
- It proposes the knowledge-circuit misalignment hypothesis, where memorized representations exist in early and late layers but are not routed into the middle reasoning layers, causing failure.
- The paper validates this across six models, from Qwen-2.5 1.5B to 7B and LLaMA-3.2 1B to 8B, and across two domains, biomedical and academic.
- A fixed heuristic using only two layer-pair settings, 0.8L to 0.5L and 0.1L to 0.5L, recovers 58 to 75 percent of the oracle improvement.
- Self-patching is far better than CoT prompting or irrelevant patching, which supports the claim that the root cause is structural.
Paper links
External research summaries. These are not HDATF publications or measured product results.