Learning to Reason in 13 Parameters

Published
Source
arXiv
Paper number
114
Field
Fine-tuning / Efficiency
arXiv ID
2602.04118

Key points

  • Fine-tuning large language models traditionally requires updating billions of parameters, which leads to high compute and memory costs.
  • Existing parameter-efficient fine-tuning (PEFT) methods reduce parameter count but still operate at the scale of millions or tens of thousands, limiting their use in constrained environments.
  • It is especially important to understand the fundamental minimum number of parameters needed for effective adaptation on complex reasoning tasks.
  • The paper introduces TinyLoRA, a new PEFT method that extends LoRA-XS by replacing trainable r×r matrices with low-dimensional trainable vectors projected through fixed random tensors, greatly reducing per-module parameters.
  • It implements an aggressive weight-sharing strategy so that trainable vectors are shared across many or all layers of the model, reducing the total number of trainable parameters to single digits.
  • It uses reinforcement learning, especially GRPO, as the core training algorithm to show that sparse reward-based signals can be exploited with highly constrained parameter updates.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)