llama-3.3-70b-versatile
groq
Meta's Llama 3.3 70B — latest iteration with improved instruction following, served on Groq LPU.
- Context window
- 131,072 tokens
- Input cost
- $0.59 / 1M
- Output cost
- $0.79 / 1M
- Latency (p50)
- 202 ms
Loading...
Run side-by-side checks for pricing, context window, and latency.
Price (lower is better), Speed/Latency (lower is better), Performance (higher is better)
Qwen: Qwen2.5 VL 32B Instruct is 195% cheaper
💰Qwen: Qwen2.5 VL 32B Instruct is 66% cheaper than llama-3.3-70b-versatile
⚡Qwen: Qwen2.5 VL 32B Instruct is 202ms faster than llama-3.3-70b-versatile
ilsp
Open-source CoRM model from ilsp — available for download and self-hosting on Hugging Face.
dr-housemd
Open-source GLM-4.7-Flash-exl3-3bpw-H4 model from dr-housemd — available for download and self-hosting on Hugging Face.
redhatai
Open-source gpt-oss-120b-essential model from redhatai — available for download and self-hosting on Hugging Face.
Let our experts help you evaluate models based on your actual use case and budget. We'll provide a free analysis and actionable recommendations.
groq
Meta's Llama 3.3 70B — latest iteration with improved instruction following, served on Groq LPU.
qwen
Qwen2.5-VL-32B is a multimodal vision-language model fine-tuned through reinforcement learning for enhanced mathematical reasoning, structured outputs, and visual problem-solving capabilities. It excels at visual analysis tasks, including object recognition, textual...