Tina: Tiny Reasoning Models via LoRA

Published
Source
arXiv
Paper number
065
Field
Reasoning / Small Models
arXiv ID
2504.15777

Key points

  • Achieving robust multi-step reasoning in large language models remains a major challenge.
  • Current LLM reasoning enhancement methods, such as supervised fine-tuning and reinforcement learning, are resource-intensive and often require substantial compute and large models.
  • The high cost and complexity of RL pipelines limit both accessibility and broad research on RL-based reasoning for language models.
  • It uses the small 1.5-billion-parameter base model DeepSeek-R1-Distill-Qwen-1.5B, which is known for its initial reasoning aptitude.
  • It applies Low-Rank Adaptation (LoRA) for parameter-efficient reinforcement learning with a GRPO-style algorithm, updating only a subset of the model parameters.
  • To ensure cost efficiency and direct comparability with existing models, it uses a minimalist training setup with two NVIDIA L40S GPUs and public reasoning datasets.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)