Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

Published
Source
arXiv
Paper number
181
Field
Agents / Coding / RAG
arXiv ID
2605.15184

Key points

  • Prior work often evaluates retrieval strategies for LLM agents in isolation rather than within a full iterative agent workflow.
  • The interaction between retrieval strategies, agent structure, custom versus provider-native CLI, and tool-calling paradigms, inline versus file-based result passing, has remained largely unexplored.
  • Understanding how retrieval strategies behave as corpus noise increases has been limited, even though this is an important aspect of real-world agent deployments.
  • We conducted a multifaceted empirical study comparing lexical retrieval, grep, and semantic retrieval, vector search, across different agent harnesses and tool-calling structures.
  • The evaluation used a subset of 116 questions from the LongMemEval benchmark and focused on multi-session conversations with varying levels of irrelevant distractor content.
  • We evaluated five different LLMs and used an auxiliary LLM judge to ensure consistent and objective agent accuracy across all experimental conditions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)