AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Published
Source
arXiv
Paper number
562
Field
AI / General
arXiv ID
2607.02255

Key points

  • The key idea is that memory is not storage but a decision visibility contract, so per-decision state is composed through typed retrieval instead of raw transcript accumulation.
  • Five slots, L1 through L5, can each be ablated independently, and the L5 trigger skill contributes the most, moving performance from 3 out of 10 to 6 out of 10 with directional p around 0.37.
  • The prompt stays around 5K tokens, while STS2MCP expands to about 527K tokens per call in a single run, which is an order-of-magnitude to two-order-of-magnitude gain in token efficiency.
  • It uses Slay the Spire 2 A0 as the testbed, where human win rate is 16% and frontier LLMs score zero, so the task is hard but not saturated.
  • It releases 298 trajectories, SHA-anchored L4 and L5 snapshots, prompt logs, and Wilson and bootstrap scripts.
  • A cross-backbone probe across Gemini, Qwen, and DeepSeek verifies the same ablation surface.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)