Cheapest long context models
Models with at least 128,000 tokens of context, ranked by input rate. A note appears when the catalog publishes a higher long context tier. Example request stays 1000 input and 500 output, so it does not cross those tiers.
All pricing ranks · Cheapest LLM API · Cheapest frontier · Cheapest long replies · Long context · Price one prompt · Cost calculator
Ranked by published rates
| # | Model | Provider | Context | Input / 1M | Output / 1M | Example | 10k requests | |
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 1.5 Flash-8B | 1M | $0.037/1M | $0.150/1M | $0.000112 | $1.12 | Cost | |
| 2 | Command R7B | Cohere | 128K | $0.037/1M | $0.150/1M | $0.000112 | $1.12 | Cost |
| 3 | Llama 3.2 1B (Groq) | Groq | 128K | $0.040/1M | $0.040/1M | $0.00006 | $0.6 | Cost |
| 4 | Llama 3.1 8B Instant (Groq) | Groq | 128K | $0.050/1M | $0.080/1M | $0.00009 | $0.9 | Cost |
| 5 | GPT-5 Nano | OpenAI | 128K | $0.050/1M | $0.400/1M | $0.00025 | $2.5 | Cost |
| 6 | Llama 3.2 3B (Groq) | Groq | 128K | $0.060/1M | $0.060/1M | $0.00009 | $0.9 | Cost |
| 7 | Llama 3.2 3B (Together) | Together AI | 128K | $0.060/1M | $0.060/1M | $0.00009 | $0.9 | Cost |
| 8 | Mistral Small | Mistral | 128K | $0.060/1M | $0.180/1M | $0.00015 | $1.5 | Cost |
| 9 | Gemini 1.5 Flash | 1M | $0.075/1M | $0.300/1M | $0.000225 | $2.25 | Cost | |
| 10 | Gemini 2.0 Flash-Lite | 1M | $0.075/1M | $0.300/1M | $0.000225 | $2.25 | Cost | |
| 11 | Llama 3.2 3B (Fireworks) | Fireworks | 128K | $0.100/1M | $0.100/1M | $0.00015 | $1.5 | Cost |
| 12 | GPT-4.1 Nano | OpenAI | 1M | $0.100/1M | $0.400/1M | $0.0003 | $3.00 | Cost |
| 13 | Gemini 2.0 Flash | 1M | $0.100/1M | $0.400/1M | $0.0003 | $3.00 | Cost | |
| 14 | Gemini 2.5 Flash-Lite | 1M | $0.100/1M | $0.400/1M | $0.0003 | $3.00 | Cost | |
| 15 | Llama 4 Scout (Groq) | Groq | 128K | $0.110/1M | $0.340/1M | $0.00028 | $2.8 | Cost |
| 16 | DeepSeek Coder | DeepSeek | 128K | $0.140/1M | $0.280/1M | $0.00028 | $2.8 | Cost |
| 17 | DeepSeek V4 Flash | DeepSeek | 1M | $0.140/1M | $0.280/1M | $0.00028 | $2.8 | Cost |
| 18 | Ministral 8B | Mistral | 128K | $0.150/1M | $0.150/1M | $0.000225 | $2.25 | Cost |
| 19 | Mistral Nemo | Mistral | 128K | $0.150/1M | $0.150/1M | $0.000225 | $2.25 | Cost |
| 20 | Pixtral 12B | Mistral | 128K | $0.150/1M | $0.150/1M | $0.000225 | $2.25 | Cost |
| 21 | GPT-4o mini | OpenAI | 128K | $0.150/1M | $0.600/1M | $0.00045 | $4.5 | Cost |
| 22 | Command R | Cohere | 128K | $0.150/1M | $0.600/1M | $0.00045 | $4.5 | Cost |
| 23 | Command R (08-2024) | Cohere | 128K | $0.150/1M | $0.600/1M | $0.00045 | $4.5 | Cost |
| 24 | Llama 3.2 11B Vision (Groq) | Groq | 128K | $0.180/1M | $0.180/1M | $0.00027 | $2.7 | Cost |
| 25 | Llama 3.1 8B (Together) | Together AI | 128K | $0.180/1M | $0.180/1M | $0.00027 | $2.7 | Cost |
| 26 | Llama 3.1 8B (Fireworks) | Fireworks | 128K | $0.200/1M | $0.200/1M | $0.0003 | $3.00 | Cost |
| 27 | Gemma 3 27B | 128K | $0.200/1M | $0.400/1M | $0.0004 | $4.00 | Cost | |
| 28 | Grok 2 Mini | xAI | 131K | $0.200/1M | $0.500/1M | $0.00045 | $4.5 | Cost |
| 29 | Grok 4 Fast | xAI | 256K | $0.200/1M | $0.500/1M | $0.00045 | $4.5 | Cost |
| 30 | GPT-5.6 Luna | OpenAI | 1.1M | $0.200/1M | $1.20/1M | $0.0008 | $8.00 | Cost |
| 31 | Llama 4 Maverick (Fireworks) | Fireworks | 1M | $0.220/1M | $0.880/1M | $0.00066 | $6.6 | Cost |
| 32 | Claude 3 Haiku | Anthropic | 200K | $0.250/1M | $1.25/1M | $0.000875 | $8.75 | Cost |
| 33 | Gemini 3.1 Flash-Lite | 1M | $0.250/1M | $1.50/1M | $0.001 | $10.00 | Cost | |
| 34 | GPT-5 Mini | OpenAI | 128K | $0.250/1M | $2.00/1M | $0.00125 | $12.5 | Cost |
| 35 | Llama 4 Maverick (Together) | Together AI | 1M | $0.270/1M | $0.850/1M | $0.000695 | $6.95 | Cost |
| 36 | DeepSeek V3 | DeepSeek | 128K | $0.270/1M | $1.10/1M | $0.00082 | $8.2 | Cost |
| 37 | DeepSeek Chat | DeepSeek | 128K | $0.280/1M | $0.420/1M | $0.00049 | $4.9 | Cost |
| 38 | QwQ 32B (Groq) | Groq | 128K | $0.290/1M | $0.390/1M | $0.000485 | $4.85 | Cost |
| 39 | Qwen3 32B (Groq) | Groq | 131K | $0.290/1M | $0.590/1M | $0.000585 | $5.85 | Cost |
| 40 | Grok 3 Mini | xAI | 131K | $0.300/1M | $0.500/1M | $0.00055 | $5.5 | Cost |
| 41 | DeepSeek R1 Distill Qwen 32B | DeepSeek | 128K | $0.300/1M | $0.600/1M | $0.0006 | $6.00 | Cost |
| 42 | Gemini 2.5 Flash | 1M | $0.300/1M | $2.50/1M | $0.00155 | $15.5 | Cost | |
| 43 | GPT-4.1 Mini | OpenAI | 1M | $0.400/1M | $1.60/1M | $0.0012 | $12.00 | Cost |
| 44 | Mistral Medium 3 | Mistral | 128K | $0.400/1M | $2.00/1M | $0.0014 | $14.00 | Cost |
| 45 | DeepSeek V4 Pro | DeepSeek | 1M | $0.435/1M | $0.870/1M | $0.00087 | $8.7 | Cost |
| 46 | Mistral Large 3 | Mistral | 128K | $0.500/1M | $1.50/1M | $0.00125 | $12.5 | Cost |
| 47 | Gemini 3 Flash | 1M | $0.500/1M | $3.00/1M | $0.002 | $20.00 | Cost | |
| 48 | DeepSeek R1 | DeepSeek | 128K | $0.550/1M | $2.19/1M | $0.001645 | $16.45 | Cost |
| 49 | DeepSeek Reasoner | DeepSeek | 128K | $0.550/1M | $2.19/1M | $0.001645 | $16.45 | Cost |
| 50 | Llama 3.3 70B (Groq) | Groq | 128K | $0.590/1M | $0.790/1M | $0.000985 | $9.85 | Cost |
| 51 | DeepSeek R1 Distill Llama 70B (Groq) | Groq | 128K | $0.750/1M | $0.990/1M | $0.001245 | $12.45 | Cost |
| 52 | Claude 3.5 Haiku | Anthropic | 200K | $0.800/1M | $4.00/1M | $0.0028 | $28.00 | Cost |
| 53 | Llama 3.1 70B (Together) | Together AI | 128K | $0.880/1M | $0.880/1M | $0.00132 | $13.2 | Cost |
| 54 | Llama 3.3 70B (Together) | Together AI | 128K | $0.880/1M | $0.880/1M | $0.00132 | $13.2 | Cost |
| 55 | DeepSeek R1 (Fireworks) | Fireworks | 128K | $0.900/1M | $0.900/1M | $0.00135 | $13.5 | Cost |
| 56 | DeepSeek V3 (Fireworks) | Fireworks | 128K | $0.900/1M | $0.900/1M | $0.00135 | $13.5 | Cost |
| 57 | Llama 3.1 70B (Fireworks) | Fireworks | 128K | $0.900/1M | $0.900/1M | $0.00135 | $13.5 | Cost |
| 58 | Llama 3.3 70B (Fireworks) | Fireworks | 128K | $0.900/1M | $0.900/1M | $0.00135 | $13.5 | Cost |
| 59 | Command Nightly | Cohere | 128K | $1.00/1M | $2.00/1M | $0.002 | $20.00 | Cost |
| 60 | Codestral | Mistral | 256K | $1.00/1M | $3.00/1M | $0.0025 | $25.00 | Cost |
| 61 | Claude Haiku 4.5 | Anthropic | 200K | $1.00/1M | $5.00/1M | $0.0035 | $35.00 | Cost |
| 62 | o1-mini | OpenAI | 128K | $1.10/1M | $4.40/1M | $0.0033 | $33.00 | Cost |
| 63 | o3-mini | OpenAI | 200K | $1.10/1M | $4.40/1M | $0.0033 | $33.00 | Cost |
| 64 | o4-mini | OpenAI | 200K | $1.10/1M | $4.40/1M | $0.0033 | $33.00 | Cost |
| 65 | DeepSeek V3 (Together) | Together AI | 128K | $1.25/1M | $1.25/1M | $0.001875 | $18.75 | Cost |
| 66 | Gemini 1.5 Pro | 2M | $1.25/1M | $5.00/1M | $0.00375 | $37.5 | Cost | |
| 67 | Gemini 2.0 Pro Exp | 2M | $1.25/1M | $5.00/1M | $0.00375 | $37.5 | Cost | |
| 68 | GPT-5 | OpenAI | 400K | $1.25/1M | $10.00/1M | $0.00625 | $62.5 | Cost |
| 69 | Gemini 2.5 Pro (higher rate after 200K) | 1M | $1.25/1M | $10.00/1M | $0.00625 | $62.5 | Cost | |
| 70 | Grok 4.6 (higher rate after 200K) | xAI | 500K | $2.00/1M | $6.00/1M | $0.005 | $50.00 | Cost |
| 71 | Mistral Large (2407) | Mistral | 128K | $2.00/1M | $6.00/1M | $0.005 | $50.00 | Cost |
| 72 | GPT-4.1 | OpenAI | 1M | $2.00/1M | $8.00/1M | $0.006 | $60.00 | Cost |
| 73 | GPT-4.1 (2025-04-14) | OpenAI | 1M | $2.00/1M | $8.00/1M | $0.006 | $60.00 | Cost |
| 74 | o3 | OpenAI | 200K | $2.00/1M | $8.00/1M | $0.006 | $60.00 | Cost |
| 75 | Claude Sonnet 5 | Anthropic | 1M | $2.00/1M | $10.00/1M | $0.007 | $70.00 | Cost |
| 76 | Grok 2 | xAI | 131K | $2.00/1M | $10.00/1M | $0.007 | $70.00 | Cost |
| 77 | GPT-5.6 Terra | OpenAI | 1.1M | $2.00/1M | $12.00/1M | $0.008 | $80.00 | Cost |
| 78 | Gemini 3.1 Pro (higher rate after 200K) | 1M | $2.00/1M | $12.00/1M | $0.008 | $80.00 | Cost | |
| 79 | GPT-4o | OpenAI | 128K | $2.50/1M | $10.00/1M | $0.0075 | $75.00 | Cost |
| 80 | GPT-4o (2024-08-06) | OpenAI | 128K | $2.50/1M | $10.00/1M | $0.0075 | $75.00 | Cost |
| 81 | Command A | Cohere | 256K | $2.50/1M | $10.00/1M | $0.0075 | $75.00 | Cost |
| 82 | Command R+ | Cohere | 128K | $2.50/1M | $10.00/1M | $0.0075 | $75.00 | Cost |
| 83 | Command R+ (08-2024) | Cohere | 128K | $2.50/1M | $10.00/1M | $0.0075 | $75.00 | Cost |
| 84 | DeepSeek R1 (Together) | Together AI | 128K | $3.00/1M | $7.00/1M | $0.0065 | $65.00 | Cost |
| 85 | Claude 3.5 Sonnet | Anthropic | 200K | $3.00/1M | $15.00/1M | $0.0105 | $105.00 | Cost |
| 86 | Claude 3 Sonnet | Anthropic | 200K | $3.00/1M | $15.00/1M | $0.0105 | $105.00 | Cost |
| 87 | Claude Sonnet 4.6 | Anthropic | 1M | $3.00/1M | $15.00/1M | $0.0105 | $105.00 | Cost |
| 88 | Grok 3 | xAI | 131K | $3.00/1M | $15.00/1M | $0.0105 | $105.00 | Cost |
| 89 | Grok 4 | xAI | 256K | $3.00/1M | $15.00/1M | $0.0105 | $105.00 | Cost |
| 90 | ChatGPT-4o Latest | OpenAI | 128K | $5.00/1M | $15.00/1M | $0.0125 | $125.00 | Cost |
| 91 | Grok Beta | xAI | 131K | $5.00/1M | $15.00/1M | $0.0125 | $125.00 | Cost |
| 92 | Claude Opus 4.6 | Anthropic | 1M | $5.00/1M | $25.00/1M | $0.0175 | $175.00 | Cost |
| 93 | Claude Opus 4.8 | Anthropic | 1M | $5.00/1M | $25.00/1M | $0.0175 | $175.00 | Cost |
| 94 | GPT-5.6 Sol | OpenAI | 1.1M | $5.00/1M | $30.00/1M | $0.02 | $200.00 | Cost |
| 95 | Claude 2.1 | Anthropic | 200K | $8.00/1M | $24.00/1M | $0.02 | $200.00 | Cost |
| 96 | GPT-4 Turbo | OpenAI | 128K | $10.00/1M | $30.00/1M | $0.025 | $250.00 | Cost |
| 97 | o1 | OpenAI | 200K | $15.00/1M | $60.00/1M | $0.045 | $450.00 | Cost |
| 98 | o1-preview | OpenAI | 128K | $15.00/1M | $60.00/1M | $0.045 | $450.00 | Cost |
| 99 | Claude 3 Opus | Anthropic | 200K | $15.00/1M | $75.00/1M | $0.0525 | $525.00 | Cost |
Related tools
Related guides
Context windows and overflow · Will my document fit? Context window calculator · Context budgeting for RAG · What happens when you exceed the context window · Why you must reserve output tokens
Sources and references
Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.
Pricing rank FAQ
- What counts as long context here?
- A published context window of 128,000 tokens or more.
- Why mention a higher rate after a threshold?
- Some models charge more once the prompt crosses a published token line. The example request is short on purpose. Use the calculator with your real document size to see the tier.
- Is a larger window always cheaper?
- No. A wide window can still have a high per token rate, and some models raise rates on long prompts.