Loading...
Loading...
Long-context models can process entire codebases, legal contracts, research papers, or books in a single prompt — eliminating the need for chunking and retrieval pipelines. Context windows have grown from 4k to 1M tokens in under three years, and the best models maintain high accuracy throughout their full context length.
| # | Model | Input $/1M | |
|---|---|---|---|
| #1 | Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/gemini-pro-1.5),... | $0.07 | Details |
| #2 | Granite 4.0 instruct models deliver strong performance across benchmarks, achieving industry-leading results in key agentic tasks like instruction following and function calling. These efficiencies make the models well-suited for a wide range of use cases like retrieval-augmented generation (RAG), multi-agent workflows, and edge deployments. | $0.02 | Details |
Not sure which model fits your budget and latency requirements?
Use the comparison tool to evaluate any two models side-by-side, or the cost calculator to project your monthly spend.
| #3 | Qwen-Turbo, based on Qwen2.5, is a 1M context model that provides fast speed and low cost, suitable for simple tasks. | $0.03 | Details |
| #4 | Amazon Nova Micro is the fastest and most cost-effective text-only model in the Nova family, optimized for speed and low latency. Ideal for customer service, summarization, and translation at scale. | $0.04 | Use |
| #5 | OLMo-2 32B Instruct is a supervised instruction-finetuned variant of the OLMo-2 32B March 2025 base model. It excels in complex reasoning and instruction-following tasks across diverse benchmarks such as GSM8K,... | $0.05 | Details |
| #6 | Meta's Llama 3.1 8B served on Groq's LPU for ultra-low latency — ideal for fast, lightweight text tasks. | $0.05 | Use |
| #7 | Amazon Nova Lite is a very low-cost multimodal model that can process image, video, and text inputs. Fast and accurate for a wide range of tasks requiring visual and language understanding. | $0.06 | Use |
| #8 | GLM-4.7-Flash is a fast and efficient multilingual text generation model with a 131,072 token context window. Optimized for dialogue, instruction-following, and multi-turn tool calling across 100+ languages. | $0.06 | Details |
| #9 | BAAI general embedding (Base) model that transforms any given text into a 768-dimensional vector | $0.07 | Details |
| #10 | ERNIE-4.5-21B-A3B-Thinking is Baidu's upgraded lightweight MoE model, refined to boost reasoning depth and quality for top-tier performance in logical puzzles, math, science, coding, text generation, and expert-level academic benchmarks. | $0.07 | Details |