LLM Post-Training: A Deep Dive into Reasoning Large Language Models
- Published
- Source
- arXiv
- Paper number
- 040
- Field
- Reasoning / Survey
- arXiv ID
- 2502.21321
Key points
- Pretrained large language models (LLMs) often struggle with factual accuracy and produce hallucinated or logically inconsistent outputs.
- LLM reasoning is primarily based on statistical patterns rather than explicit logical inference, which limits performance on complex reasoning tasks.
- Despite massive pretraining, LLMs may still be misaligned with human intent, preferences, and ethical considerations, so additional refinement is needed for trustworthy deployment.
- This paper systematically surveys and categorizes existing post-training methodologies for LLMs, with a focus on their role in improving reasoning ability.
- It organizes these techniques into three connected stages: fine-tuning (for example, instruction, CoT, PEFT), reinforcement learning (for example, RLHF, DPO, RLAIF), and test-time scaling (for example, CoT prompting, Tree-of-Thoughts, MCTS).
- The survey highlights the specific points at which each technique helps refine LLM behavior, improve robustness, and ensure adaptability across applications.
Paper links
External research summaries. These are not HDATF publications or measured product results.