AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Published
Source
arXiv
Paper number
1076
Field
LLMs / NLP
arXiv ID
2609.08936

Key points

  • Taking natural-language instructions and audio context as common inputs, it covers speech generation, content editing, restoration and separation, and speaker-attribute or acoustic editing.
  • It combines semantic conditioning from a multimodal language model with acoustic conditioning from an audio VAE, and jointly trains generation and editing after generation-centric pretraining.
  • Different post-training signals are applied to generation and editing, and AuK-Flash reports a 4.5x speedup under matched conditions with four-step inference.
  • Handling diverse audio tasks through one unified interface makes it a foundation for building speech production and editing tools.
  • Understanding compound instructions and generalizing to new instruction combinations remain open, and the current system partly relies on task classification and instruction completion.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)