MEDEC
- Published
- Source
- arXiv
- Paper number
- 005
- Field
- Medical AI
- arXiv ID
- 2412.19260
Key points
- Instead of asking medical QA questions, MEDEC asks a model to verify the accuracy and consistency of existing or generated clinical records.
- The dataset contains 3,848 clinical texts covering diagnosis, management, treatment, pharmacotherapy, and etiology-related errors.
- The evaluation compares 17 shared-task systems and frontier LLMs, with physicians used as the human reference point.
- In conclusion, LLMs show useful error-detection ability but still fall short of physician-level performance, underscoring the need for careful deployment in clinical documentation.
Paper links
External research summaries. These are not HDATF publications or measured product results.