Loading...
Loading...
🌐 All Models
461 models · Page 1 of 13
Open-source nsfwvision-qwen3-vl-8b-v3-GGUF model from dummy9996 — available for download and self-hosting on Hugging Face.
Open-source llava-onevision-qwen2-7b-ov-hf-finetuned model from multimodal-colab — available for download and self-hosting on Hugging Face.
Open-source Qwen3-VL-8b-Instruct-finetuned model from multimodal-colab — available for download and self-hosting on Hugging Face.
Open-source Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS model from aeon-7 — available for download and self-hosting on Hugging Face.
Open-source Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP model from aeon-7 — available for download and self-hosting on Hugging Face.
Open-source Krea2-Loras model from deverstyle — available for download and self-hosting on Hugging Face.
Open-source anima-mlx model from xocialize — available for download and self-hosting on Hugging Face.
Open-source AZ model from az2671 — available for download and self-hosting on Hugging Face.
Open-source shadow-vision-7b-v2 model from lhordking — available for download and self-hosting on Hugging Face.
Open-source 100m_image_new model from eliochampaney — available for download and self-hosting on Hugging Face.
Open-source Zeta-Chroma model from lodestones — available for download and self-hosting on Hugging Face.
Open-source vision-global-tmdm-price-qlora-qwenvl3 model from acchf — available for download and self-hosting on Hugging Face.
Building upon Mistral Small 3 (2501), Mistral Small 3.1 (2503) adds state-of-the-art vision understanding and enhances long context capabilities up to 128k tokens without compromising text performance. With 24 billion parameters, this model achieves top-tier capabilities in both text and vision tasks.
The Llama 3.2-Vision instruction-tuned models are optimized for visual recognition, image reasoning, captioning, and answering general questions about an image.
UForm-Gen is a small generative vision-language model primarily designed for Image Captioning and Visual Question Answering. The model was pre-trained on the internal image captioning dataset and fine-tuned on public instructions datasets: SVIT, LVIS, VQAs datasets.
FLUX.2 [dev] is an image model from Black Forest Labs where you can generate highly realistic and detailed images, with multi-reference support.
Stable Diffusion is a latent text-to-image diffusion model capable of generating photo-realistic images. Img2img generate a new image from an input image with Stable Diffusion.
Lucid Origin from Leonardo.AI is their most adaptable and prompt-responsive model to date. Whether you're generating images with sharp graphic design, stunning full-HD renders, or highly specific creative direction, it adheres closely to your prompts, renders text with accuracy, and supports a wide array of visual styles and aesthetics – from stylized concept art to crisp product mockups.
Meta's Llama 4 Scout is a 17 billion parameter model with 16 experts that is natively multimodal. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding.
Gemma 3 models are well-suited for a variety of text generation and image understanding tasks, including question answering, summarization, and reasoning. Gemma 3 models are multimodal, handling text and image input and generating text output, with a large, 128K context window, multilingual support in over 140 languages, and is available in more sizes than previous versions.
FLUX.2 [klein] 9B is a 9 billion parameter model that can generate images from text descriptions and supports multi-reference editing capabilities.
Kimi K2.5 is a frontier-scale open-source model with a 256k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
LLaVA is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal instruction-following data. It is an auto-regressive language model, based on the transformer architecture.
FLUX.1 [schnell] is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions.
SDXL-Lightning is a lightning-fast text-to-image generation model. It can generate high-quality 1024px images in a few steps.
Stable Diffusion model that has been fine-tuned to be better at photorealism without sacrificing range.
Phoenix 1.0 is a model by Leonardo.Ai that generates images with exceptional prompt adherence and coherent text.
Diffusion-based text-to-image generative model by Stability AI. Generates and modify images based on text prompts.
FLUX.2 [klein] is an ultra-fast, distilled image model. It unifies image generation and editing in a single model, delivering state-of-the-art quality enabling interactive workflows, real-time previews, and latency-critical applications.
Kimi K2.6 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
Stable Diffusion Inpainting is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input, with the extra capability of inpainting the pictures by using a mask.
Claude 3.7 Sonnet — available via AWS Bedrock (us-east-1).
Claude Sonnet 4.5 — available via AWS Bedrock (us-east-1).
Nova Canvas — available via AWS Bedrock (us-east-1).
Nova Reel — available via AWS Bedrock (us-east-1).
Gemma 3 4B IT — available via AWS Bedrock (us-east-1).