ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents

Published
Source
arXiv
Paper number
293
Field
AI / General
arXiv ID
2606.02568

Key points

  • Static benchmarks cannot validate this, and existing interactive medical benchmarks each compromise on at least one dimension.
  • This paper presents ClinEnv, an interactive benchmark that evaluates LLMs as attending physicians on real inpatient cases under a paradigm called Longitudinal Inpatient Simulation.
  • Each case is automatically organized into a sequence of ordered decision steps, and at every step the model must actively query four specialized agents before finalizing medications, procedures, or diagnoses.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)