Loading...
Loading...
Side-by-side comparison of API pricing, context window, latency, and benchmark performance. Data refreshed daily from provider APIs and ModelStop's own latency probes.
💰Qwen: Qwen2.5 Coder 7B Instruct is 94% cheaper than llama-guard-3-8b
⚡Qwen: Qwen2.5 Coder 7B Instruct is 170ms faster than llama-guard-3-8b
Price (lower is better), Speed/Latency (lower is better), Performance (higher is better)
Qwen: Qwen2.5 Coder 7B Instruct is 1513% cheaper
| Spec | llama-guard-3-8b | Qwen: Qwen2.5 Coder 7B Instruct |
|---|---|---|
| Provider | meta | qwen |
| Input price / 1M tokens | $0.48 | $0.03 |
| Output price / 1M tokens | $0.03 | $0.09 |
| Context window | 131k tokens | 33k tokens |
| Latency (p50) | 170ms | — |
| Best benchmark score | — | — |
| Open source | No | No |
Llama Guard 3 is a Llama-3.1-8B pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM inputs (prompt classification) and in LLM responses (response classification). It acts as an LLM – it generates text in its output that indicates whether a given prompt or response is safe or unsafe, and if unsafe, it also lists the content categories violated.
Qwen2.5-Coder-7B-Instruct is a 7B parameter instruction-tuned language model optimized for code-related tasks such as code generation, reasoning, and bug fixing. Based on the Qwen2.5 architecture, it incorporates enhancements like RoPE,...
Want to swap in a different model or estimate monthly costs?