Add support for https://huggingface.co/tencent/AuK-Flash
Author: AshDCreated Sep 12, 2026Updated Sep 12, 2026
https://auk-project.github.io/
AuK is a 1.5B-parameter foundational model that answers a natural-language instruction with edited or generated audio. A multimodal language model reads the instruction alongside optional reference audio; a 50 Hz VAE supplies acoustic latents; a hybrid transformer fuses both streams under a flow-matching objective.
Source: jamiepine/voicebox