Cloudflare Meta Llama 3.1 8B Instruct
Unknown provider
Meta Llama 3.1 8B Instruct via Cloudflare
- Context window
- 8,192 tokens
- Input cost
- —
- Output cost
- —
- Latency (p50)
- 335 ms
Loading...
Run side-by-side checks for pricing, context window, and latency.
Price (lower is better), Speed/Latency (lower is better), Performance (higher is better)
💰Cloudflare Meta Llama 3.1 8B Instruct is 100% cheaper than Meituan: LongCat Flash Chat
⚡Meituan: LongCat Flash Chat is 335ms faster than Cloudflare Meta Llama 3.1 8B Instruct
madras1
Open-source Yara-50M-PTBR-QA model from madras1 — available for download and self-hosting on Hugging Face.
manuel121357
Open-source Dolphin3.0-Llama3.1-8B-q4f16_1-MLC model from manuel121357 — available for download and self-hosting on Hugging Face.
taniota
Open-source Warlock-1.7B model from taniota — available for download and self-hosting on Hugging Face.
Let our experts help you evaluate models based on your actual use case and budget. We'll provide a free analysis and actionable recommendations.
Unknown provider
Meta Llama 3.1 8B Instruct via Cloudflare
meituan
LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input. It introduces a shortcut-connected MoE design to reduce...