Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes

Published
Source
arXiv
Paper number
179
Field
Agents / Security / Skills
arXiv ID
2605.00424

Key points

  • Agent skills, structured packages of instructions, scripts, and reference material that augment an LLM without modifying the model itself, have moved from convenience features to first-class deployment artifacts.
  • The runtimes that load them inherit the same problem faced by package managers and operating systems: when content claims to do something, the runtime must decide whether to trust it.
  • The paper makes its core point explicit from the start: a skill is untrusted code until it is verified, and the runtime that loads it should enforce that default rather than infer trust from signatures, permission levels, or source registries.
  • Without skill verification, a human-in-the-loop (HITL) gate must trigger on every irreversible call, which is operationally unsustainable and degrades into formal approval at nontrivial scale.
  • Treating skill verification as a separate gated procedure means HITL is invoked only for unverified artifacts, which makes the system sustainable.
  • The authors present a portable runtime profile that includes a trust schema with an explicit verification level in every skill manifest, a capability gate whose HITL policy is determined as a function of that verification level, a biconditional correctness criterion that any candidate verification procedure must satisfy under adversarial ensemble tests, and ten normative guidelines abstracted from a working open-source reference implementation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)