Tina: Tiny Reasoning Models via LoRA
- Published
- Source
- arXiv
- Paper number
- 065
- Field
- Reasoning / Small Models
- arXiv ID
- 2504.15777
Key points
- Achieving robust multi-step reasoning in large language models remains a major challenge.
- Current LLM reasoning enhancement methods, such as supervised fine-tuning and reinforcement learning, are resource-intensive and often require substantial compute and large models.
- The high cost and complexity of RL pipelines limit both accessibility and broad research on RL-based reasoning for language models.
- It uses the small 1.5-billion-parameter base model DeepSeek-R1-Distill-Qwen-1.5B, which is known for its initial reasoning aptitude.
- It applies Low-Rank Adaptation (LoRA) for parameter-efficient reinforcement learning with a GRPO-style algorithm, updating only a subset of the model parameters.
- To ensure cost efficiency and direct comparability with existing models, it uses a minimalist training setup with two NVIDIA L40S GPUs and public reasoning datasets.
Paper links
External research summaries. These are not HDATF publications or measured product results.