AI Model Catalogue
Browse 194 models across providers, modalities, and use cases.
🌐 All Models
194 models · Page 6 of 6
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
textvisionmultimodal
Run locally
Explore specs and pricing
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
textvisionmultimodal
Run locally
Explore specs and pricing
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
textvisionmultimodal
Run locally
Input$0.2000/1M
Output$1.2500/1M
Explore specs and pricing
MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step...
textvisionmultimodal
Input$0.4000/1M
Output$2.0000/1M
Explore specs and pricing
Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,...
textvisionimage
Run locally
Input$0.1000/1M
Output$0.1000/1M
Explore specs and pricing
30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...
textvisionimage
Run locally
InputFree
Output$0.0000/1M
Explore specs and pricing
Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...
textvisionimage
Run locally
InputFree
Output$0.0000/1M
Explore specs and pricing
Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...
textvisionmultimodal
Run locally
Input$2.0000/1M
Output$6.0000/1M
Explore specs and pricing
Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...
textvisionmultimodal
Run locally
Input$2.0000/1M
Output$6.0000/1M
Explore specs and pricing
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
textvisionmultimodal
Run locally
Explore specs and pricing
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
textvisionmultimodal
Run locally
Input$0.3250/1M
Output$1.9500/1M
📏1000kcontext
⭐1218.0%score
Explore specs and pricing
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
textvisionmultimodal
Run locally
Input$0.1400/1M
Output$0.4000/1M
Explore specs and pricing
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
textvisionmultimodal
Run locally
Input$0.1200/1M
Output$0.4000/1M
Explore specs and pricing
Fast-mode variant of [Opus 4.6](/anthropic/claude-opus-4.6) - identical capabilities with higher output speed at premium 6x pricing.
Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode
textvisionmultimodal
Run locally
Input$30.0000/1M
Output$150.0000/1M
Explore specs and pricing