Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes
- Published
- Source
- arXiv
- Paper number
- 179
- Field
- Agents / Security / Skills
- arXiv ID
- 2605.00424
Key points
- Agent skills, structured packages of instructions, scripts, and reference material that augment an LLM without modifying the model itself, have moved from convenience features to first-class deployment artifacts.
- The runtimes that load them inherit the same problem faced by package managers and operating systems: when content claims to do something, the runtime must decide whether to trust it.
- The paper makes its core point explicit from the start: a skill is untrusted code until it is verified, and the runtime that loads it should enforce that default rather than infer trust from signatures, permission levels, or source registries.
- Without skill verification, a human-in-the-loop (HITL) gate must trigger on every irreversible call, which is operationally unsustainable and degrades into formal approval at nontrivial scale.
- Treating skill verification as a separate gated procedure means HITL is invoked only for unverified artifacts, which makes the system sustainable.
- The authors present a portable runtime profile that includes a trust schema with an explicit verification level in every skill manifest, a capability gate whose HITL policy is determined as a function of that verification level, a biconditional correctness criterion that any candidate verification procedure must satisfy under adversarial ensemble tests, and ten normative guidelines abstracted from a working open-source reference implementation.
Paper links
External research summaries. These are not HDATF publications or measured product results.