Loading...
Loading...
Customer support is one of the highest-volume AI deployments, so cost-per-token and latency matter as much as raw quality. The best models for support are strong instruction followers that stay on-topic, handle multi-turn conversations gracefully, and keep costs predictable at scale.
| # | Model | Input $/1M | |
|---|---|---|---|
| #1 | Qwen2.5-Coder-7B-Instruct is a 7B parameter instruction-tuned language model optimized for code-related tasks such as code generation, reasoning, and bug fixing. Based on the Qwen2.5 architecture, it incorporates enhancements like RoPE,... | $0.03 | Details |
| #2 | Amazon Nova Micro is the fastest and most cost-effective text-only model in the Nova family, optimized for speed and low latency. Ideal for customer service, summarization, and translation at scale. | $0.04 | Use |
Not sure which model fits your budget and latency requirements?
Use the comparison tool to evaluate any two models side-by-side, or the cost calculator to project your monthly spend.
| #3 |
Amazon Nova Lite is a very low-cost multimodal model that can process image, video, and text inputs. Fast and accurate for a wide range of tasks requiring visual and language understanding. |
| $0.06 |
| Use |
| #4 | Amazon Titan Text Express is a generative LLM for summarization, text generation, classification, open-ended Q&A, and information extraction. Optimized for enterprise workloads via AWS Bedrock. | $0.20 | Use |
| #5 | AI21 Jamba 1.6 Mini is a lightweight Mamba-Transformer hybrid optimized for cost-effective, high-throughput inference with an impressive 256K context window. An excellent choice for document-heavy workloads on a budget. | $0.20 | Use |
| #6 | LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input. It introduces a shortcut-connected MoE design to reduce... | $0.20 | Details |
| #7 | Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude... | $0.25 | Details |
| #8 | Mercury Coder is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like Claude 3.5 Haiku... | $0.25 | Details |
| #9 | Amazon Nova Pro is a highly capable multimodal model with the best combination of accuracy, speed, and cost across a wide range of tasks. Supports text, image, and video inputs. | $0.80 | Use |
| #10 | Llemma 7B is a language model for mathematics. It was initialized with Code Llama 7B weights, and trained on the Proof-Pile-2 for 200B tokens. Llemma models are particularly strong at... | $0.80 | Details |