LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Published
Source
arXiv
Paper number
040
Field
Reasoning / Survey
arXiv ID
2502.21321

Key points

  • Pretrained large language models (LLMs) often struggle with factual accuracy and produce hallucinated or logically inconsistent outputs.
  • LLM reasoning is primarily based on statistical patterns rather than explicit logical inference, which limits performance on complex reasoning tasks.
  • Despite massive pretraining, LLMs may still be misaligned with human intent, preferences, and ethical considerations, so additional refinement is needed for trustworthy deployment.
  • This paper systematically surveys and categorizes existing post-training methodologies for LLMs, with a focus on their role in improving reasoning ability.
  • It organizes these techniques into three connected stages: fine-tuning (for example, instruction, CoT, PEFT), reinforcement learning (for example, RLHF, DPO, RLAIF), and test-time scaling (for example, CoT prompting, Tree-of-Thoughts, MCTS).
  • The survey highlights the specific points at which each technique helps refine LLM behavior, improve robustness, and ensure adaptability across applications.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)