Loading...
Browse 129 models across providers, modalities, and use cases.
🎙️ Audio & Speech
129 models · Page 3 of 4
A small audio understanding model released in July 2025
SAM-Audio is a foundation model for isolating any sound in audio using text
Image-to-video generation with optional audio, multi-shot narrative support, and faster inference
A foundation model for isolating any sound in audio using text, visual, or temporal prompts
Minimax Speech 2.8 HD focuses on high-fidelity audio generation with features like studio-grade quality, flexible emotion control, multilingual support, and voice cloning capabilities
Music generation
Kling Video 3.0: Generate cinematic videos up to 15 seconds with multi-shot control, native audio, and improved consistency
Ace Step 1.5 open source music generation model
An extension of AiCoverGen, which provides several new features and improvements, enabling users to generate audio-related content using RVC with ease. Ideal for people who want to incorporate singing functionality into their AI assistant/chatbot/vtuber,
High-fidelity video generation with portrait support, audio-to-video, retake, and extend. Text, image, and audio-driven creation up to 4K at 50 FPS.
Lightning-fast video generation with portrait support, camera controls, and synchronized audio. Up to 20 seconds at 1080p, 4K at 50 FPS.
Kling Video 3.0 Omni: Unified multimodal video generation with reference images, video editing, native audio, and multi-shot control
HeartMuLa: A Family of Open Sourced Music Foundation Models
Generate videos from images, with support for first-and-last-frame control, clip continuation, and audio synchronization using Alibaba's Wan 2.7 model
ByteDance's multimodal video generation model with native audio, multimodal reference inputs, and intelligent duration control.
Fast video generation with built-in draft mode for rapid creative iteration. Text-to-video, image-to-video, and audio-to-video in a single endpoint.
Fast video generation with text-to-video, image-to-video, and start-end-to-video modes. Up to 16 seconds at 1080p with synchronized audio.
Lo-fi hip-hop music generation with ACE-Step 1.5 + LoRA
Generate full-length songs or instrumentals from a text prompt, with optional auto-generated lyrics
Google's cost-efficient video generation model with native audio, optimized for high-volume applications
Reimagine any song in a different style — change voice, instruments, genre, and arrangement while keeping the original melody
High-fidelity video generation with text-to-video, image-to-video, and start-end-to-video modes. Up to 16 seconds at 1080p with synchronized audio.
A faster variant of Seedance 2.0 for quicker video generation with multimodal inputs and native audio.
Create a dotted waveform video from an audio file
New and improved version of Veo 3 Fast, with higher-fidelity video, context-aware audio and last frame support
New and improved version of Veo 3, with higher-fidelity video, context-aware audio, reference image and last frame support
Generate videos with audio from text prompts using Alibaba's Wan 2.7 model. 1080p, up to 15 seconds, with audio synchronization.
Generate full-length songs with vocals, lyrics, and rich instrumentation from a text prompt
zai-org/GLM-ASR-Nano-2512 is a automatic speech recognition model on Hugging Face with ~160,973 monthly downloads. Open access.