Loading...
Browse 20 models across providers, modalities, and use cases.
🎙️ Audio & Speech
20 models · Page 1 of 1
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalize to many datasets and domains without the need for fine-tuning. This is the English-only version of the Whisper Tiny model which was trained on the task of speech recognition.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation.
Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification.
Open-source whisper-medium model from openai — available for download and self-hosting on Hugging Face.
Open-source whisper-tiny model from openai — available for download and self-hosting on Hugging Face.
Open-source whisper-large-v3 model from openai — available for download and self-hosting on Hugging Face.
Open-source whisper-large-v3-turbo model from openai — available for download and self-hosting on Hugging Face.
Open-source whisper-base model from openai — available for download and self-hosting on Hugging Face.
Open-source whisper-small model from openai — available for download and self-hosting on Hugging Face.
The gpt-4o-audio-preview model adds support for audio inputs as prompts. This enhancement allows the model to detect nuances within audio recordings and add depth to generated user experiences. Audio outputs...
A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...