Agentic-imodels: Evolving agentic interpretability tools via autoresearch

Published
Source
arXiv
Paper number
178
Field
Agents / Interpretability / AutoML
arXiv ID
2605.03808

Key points

  • Existing interpretability machine learning tools are designed for human interpretation and therefore produce outputs, such as visualizations, that AI agents cannot reliably parse or use.
  • This mismatch leads to unreliable analysis, obscured reasoning, and inefficient workflows in autonomous data science systems.
  • The lack of automated, model-agnostic metrics for agent interpretability has slowed the development of AI-centered interpretability tools.
  • An agentic self-research loop repeatedly modifies Python model implementations based on prediction performance and feedback from a new LLM-scored interpretability score.
  • It introduces a new agent interpretability score that quantifies an LLM's ability to reason about a model from its string representation using 200 separate automated tests.
  • The framework optimizes models toward the joint goal of high predictive performance, measured by ranked RMSE across 65 datasets, and high agent interpretability, measured by LLM test pass rate.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)