Loading...
Browse 132 models across providers, modalities, and use cases.
🎙️ Audio & Speech
132 models · Page 3 of 4
A small audio understanding model released in July 2025
A mini audio understanding model released in July 2025
SAM-Audio is a foundation model for isolating any sound in audio using text
Image-to-video generation with optional audio, multi-shot narrative support, and faster inference
A foundation model for isolating any sound in audio using text, visual, or temporal prompts
Minimax Speech 2.8 HD focuses on high-fidelity audio generation with features like studio-grade quality, flexible emotion control, multilingual support, and voice cloning capabilities
Music generation
Kling Video 3.0: Generate cinematic videos up to 15 seconds with multi-shot control, native audio, and improved consistency
HeartMuLa: A Family of Open Sourced Music Foundation Models
High-fidelity video generation with portrait support, audio-to-video, retake, and extend. Text, image, and audio-driven creation up to 4K at 50 FPS.
Lightning-fast video generation with portrait support, camera controls, and synchronized audio. Up to 20 seconds at 1080p, 4K at 50 FPS.
Ace Step 1.5 open source music generation model
An extension of AiCoverGen, which provides several new features and improvements, enabling users to generate audio-related content using RVC with ease. Ideal for people who want to incorporate singing functionality into their AI assistant/chatbot/vtuber,
Kling Video 3.0 Omni: Unified multimodal video generation with reference images, video editing, native audio, and multi-shot control
Lo-fi hip-hop music generation with ACE-Step 1.5 + LoRA
Generate videos with audio from text prompts using Alibaba's Wan 2.7 model. 1080p, up to 15 seconds, with audio synchronization.
New and improved version of Veo 3, with higher-fidelity video, context-aware audio, reference image and last frame support
New and improved version of Veo 3 Fast, with higher-fidelity video, context-aware audio and last frame support
Create a dotted waveform video from an audio file
A faster variant of Seedance 2.0 for quicker video generation with multimodal inputs and native audio.
High-fidelity video generation with text-to-video, image-to-video, and start-end-to-video modes. Up to 16 seconds at 1080p with synchronized audio.
Google's cost-efficient video generation model with native audio, optimized for high-volume applications
Generate full-length songs or instrumentals from a text prompt, with optional auto-generated lyrics
Fast video generation with text-to-video, image-to-video, and start-end-to-video modes. Up to 16 seconds at 1080p with synchronized audio.
Generate videos from images, with support for first-and-last-frame control, clip continuation, and audio synchronization using Alibaba's Wan 2.7 model
Fast video generation with built-in draft mode for rapid creative iteration. Text-to-video, image-to-video, and audio-to-video in a single endpoint.
Generate full-length songs with vocals, lyrics, and rich instrumentation from a text prompt