Loading...
Loading...
👁️ Vision & Multimodal
463 models · Page 6 of 13
Open-source stable-diffusion-v1-5 model from stable-diffusion-v1-5 — available for download and self-hosting on Hugging Face.
Open-source FLUX.1-dev model from black-forest-labs — available for download and self-hosting on Hugging Face.
Open-source stable-diffusion-v1-5 model from crynux-network — available for download and self-hosting on Hugging Face.
Open-source novaAnimeXL_ilV140 model from frankjoshua — available for download and self-hosting on Hugging Face.
Open-source one-obsession-17-red-sdxl model from john6666 — available for download and self-hosting on Hugging Face.
Open-source animagine-xl-4.0 model from cagliostrolab — available for download and self-hosting on Hugging Face.
Open-source diving-illustrious-real-asian-v50-sdxl model from john6666 — available for download and self-hosting on Hugging Face.
Open-source sdxl-turbo model from stabilityai — available for download and self-hosting on Hugging Face.
Open-source stable-diffusion-v1-4 model from compvis — available for download and self-hosting on Hugging Face.
Open-source Qwen-Image-Lightning model from lightx2v — available for download and self-hosting on Hugging Face.
Open-source playground-v2.5-1024px-aesthetic model from playgroundai — available for download and self-hosting on Hugging Face.
Open-source stable-diffusion-xl-base-1.0 model from stabilityai — available for download and self-hosting on Hugging Face.
Open-source sd-turbo model from stabilityai — available for download and self-hosting on Hugging Face.
Open-source Z-Image-Turbo model from tongyi-mai — available for download and self-hosting on Hugging Face.
Open-source FLUX.1-schnell model from black-forest-labs — available for download and self-hosting on Hugging Face.
Open-source sdxl-turbo model from crynux-network — available for download and self-hosting on Hugging Face.
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Cohere's latest multimodal embedding model supporting text and images for advanced semantic search.
Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Our frontier-class multimodal model released May 2025.
Agentic image model optimized for robust, high-precision generations supporting font control
FireRed-Image-Edit is a general-purpose image editing model that delivers high-fidelity and consistent editing across a wide range of scenarios.