Loading...
Browse 6 models across providers, modalities, and use cases.
👁️ Vision & Multimodal
6 models · Page 1 of 1
The Llama 3.2-Vision instruction-tuned models are optimized for visual recognition, image reasoning, captioning, and answering general questions about an image.
Meta's Llama 4 Scout is a 17 billion parameter model with 16 experts that is natively multimodal. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding.
Llama 3.2 11B Instruct — available via AWS Bedrock (us-east-1).
Llama 3.2 90B Instruct — available via AWS Bedrock (us-east-1).
Llama 4 Scout 17B Instruct — available via AWS Bedrock (us-east-1).
Llama 4 Maverick 17B Instruct — available via AWS Bedrock (us-east-1).