Llama-Guard-3-8B
meta-llama
Open-source Llama-Guard-3-8B model from meta-llama — available for download and self-hosting on Hugging Face.
- Context window
- 131,072 tokens
- Input cost
- $0.48 / 1M
- Output cost
- $0.03 / 1M
- Latency (p50)
- —
Loading...
Run side-by-side checks for pricing, context window, and latency.
Price (lower is better), Speed/Latency (lower is better), Performance (higher is better)
Meituan: LongCat Flash Chat is 140% cheaper
💰Meituan: LongCat Flash Chat is 58% cheaper than Llama-Guard-3-8B
giannisan
Open-source Laguna-S-2.1-Q2K-pulsar model from giannisan — available for download and self-hosting on Hugging Face.
vontra
Open-source Solar-Open2-250B-MLX-6bit model from vontra — available for download and self-hosting on Hugging Face.
atomicchat
Open-source ornith-35b-MLX-6bit model from atomicchat — available for download and self-hosting on Hugging Face.
Let our experts help you evaluate models based on your actual use case and budget. We'll provide a free analysis and actionable recommendations.
meta-llama
Open-source Llama-Guard-3-8B model from meta-llama — available for download and self-hosting on Hugging Face.
meituan
LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input. It introduces a shortcut-connected MoE design to reduce...