ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- Published
- Source
- arXiv
- Paper number
- 082
- Field
- Reasoning / Parallelism
- arXiv ID
- 2509.04475
Key points
- Sequential LLM reasoning shows tunnel vision, where early mistakes cause diminishing returns and stagnation even as compute increases.
- Current depth-scaling approaches become inefficient beyond a certain point and fail to unlock further gains on complex reasoning tasks.
- Existing parallel or search-based methods often depend on external verifiers or domain-specific components, or they add substantial extra compute.
- ParaThinker uses a two-stage process: parallel reasoning that generates multiple independent thought paths, followed by summarization that integrates those paths into a single final answer.
- It uses special control tokens, `think i` and `summary`, to guide the model in generating diverse reasoning paths and then merging them.
- Thought-specific positional embeddings augment standard RoPE to resolve positional ambiguity across parallel paths, ensuring distinct processing and effective summarization.
Paper links
External research summaries. These are not HDATF publications or measured product results.