HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness

Published
Source
arXiv
Paper number
170
Field
Reasoning / Agents / Coding
arXiv ID
2605.02396

Key points

  • The concrete mechanisms that drive performance inside complex LLM-agent orchestration frameworks remain largely unidentified.
  • Current LLM-agent designs often depend on broad systems engineering without a clear abstraction for the core reasoning process.
  • Existing test-time scaling, or TTS, strategies often require specialized structural changes or large amounts of post-training, so they lack a universal and portable skill representation.
  • The authors propose HEAVYSKILL, a two-stage reasoning pipeline consisting of parallel reasoning that generates K independent trajectories and sequential deliberation that processes and synthesizes those trajectories.
  • They implement a serialized memory cache that stores and organizes reasoning trajectories, which helps context management and avoids positional bias during deliberation.
  • They improve portability by distilling the focused thinking workflow into a readable skill document that any LLM orchestrator can execute autonomously from in-context instructions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)