header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Tencent Open-Sources "Audio Version of Nano Banana": Rewording and Voice Swapping All in One, Voice Editing Significantly Tops the Charts

Beating AI News Flash: Tencent Hunyuan has open-sourced AuK, a 1.5-billion-parameter speech generation and editing model, which the official team directly calls the "Nano Banana of audio." Voice cloning, word replacement, emotion switching, accent removal, speed and pitch adjustment, noise reduction, and multi-speaker and music separation are all packed into a single model, controllable through natural language.


AuK can directly edit existing recordings. If a few words are misspoken, only those words need to be changed—no need to re-record the entire segment. Without changing the content, users can also switch emotions, timbres, or accents. It also supports editing lyrics, removing laughter and coughing sounds, and converting normal speech into whispers. Voice cloning no longer requires providing a transcript for the reference recording first—just a segment of audio and new text is enough to generate output.


In official tests, AuK leads in multiple evaluations for speech generation and general speech editing. It scored 49.73 on SpeechEditBench, while the second-place Ming-UniAudio scored 28.70. The fast version, AuK-Flash, after distillation, requires only 4 inference steps and is about 4.5 times faster. However, the paper also acknowledges that AuK currently cannot stably understand all arbitrary natural language instructions and still relies on a Prompt Enhancer to first determine the task, organize parameters, and rewrite instructions. The code and model weights have already been made public under the MIT license.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish