MMAE: A Massive Multitask Audio Editing Benchmark

Published
Source
arXiv
Paper number
365
Field
LLMs / NLP
arXiv ID
2606.07229

Key points

  • It establishes a unified taxonomy for audio editing for the first time, spanning seven modalities, six complexity levels, and eight operation types.
  • It introduces an evaluation paradigm that decomposes free-form editing tasks into a verifiable multidimensional rubric, which enables objective and interpretable assessment.
  • It builds the benchmark from 2,000 high-quality samples and 17,741 evaluation rubrics selected through a human-agent collaboration pipeline.
  • Evaluation of five state-of-the-art models shows an EMR below 5 percent and 0 percent on complex mixed modalities, which exposes the fundamental limitations of current systems.
  • It confirms that introducing an external agent planner does not consistently help and that both understanding and generation remain bottlenecks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)