AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
- Published
- Source
- arXiv
- Paper number
- 562
- Field
- AI / General
- arXiv ID
- 2607.02255
Key points
- The key idea is that memory is not storage but a decision visibility contract, so per-decision state is composed through typed retrieval instead of raw transcript accumulation.
- Five slots, L1 through L5, can each be ablated independently, and the L5 trigger skill contributes the most, moving performance from 3 out of 10 to 6 out of 10 with directional p around 0.37.
- The prompt stays around 5K tokens, while STS2MCP expands to about 527K tokens per call in a single run, which is an order-of-magnitude to two-order-of-magnitude gain in token efficiency.
- It uses Slay the Spire 2 A0 as the testbed, where human win rate is 16% and frontier LLMs score zero, so the task is hard but not saturated.
- It releases 298 trajectories, SHA-anchored L4 and L5 snapshots, prompt logs, and Wilson and bootstrap scripts.
- A cross-backbone probe across Gemini, Qwen, and DeepSeek verifies the same ablation surface.
Paper links
External research summaries. These are not HDATF publications or measured product results.