Learning to Reason in 13 Parameters
- Published
- Source
- arXiv
- Paper number
- 114
- Field
- Fine-tuning / Efficiency
- arXiv ID
- 2602.04118
Key points
- Fine-tuning large language models traditionally requires updating billions of parameters, which leads to high compute and memory costs.
- Existing parameter-efficient fine-tuning (PEFT) methods reduce parameter count but still operate at the scale of millions or tens of thousands, limiting their use in constrained environments.
- It is especially important to understand the fundamental minimum number of parameters needed for effective adaptation on complex reasoning tasks.
- The paper introduces TinyLoRA, a new PEFT method that extends LoRA-XS by replacing trainable r×r matrices with low-dimensional trainable vectors projected through fixed random tensors, greatly reducing per-module parameters.
- It implements an aggressive weight-sharing strategy so that trainable vectors are shared across many or all layers of the model, reducing the total number of trainable parameters to single digits.
- It uses reinforcement learning, especially GRPO, as the core training algorithm to show that sparse reward-based signals can be exploited with highly constrained parameter updates.
Paper links
External research summaries. These are not HDATF publications or measured product results.