Loading...
Loading...
🎙️ Audio & Speech
132 models · Page 4 of 4
ByteDance's multimodal video generation model with native audio, multimodal reference inputs, and intelligent duration control.
Reimagine any song in a different style — change voice, instruments, genre, and arrangement while keeping the original melody
zai-org/GLM-ASR-Nano-2512 is a automatic speech recognition model on Hugging Face with ~160,973 monthly downloads. Open access.