Home / Pricing / Long replies

Cheapest models for long replies

Ranked by output rate per 1M tokens. The example uses 1000 input tokens and 2000 output tokens, which is closer to support chat and writing than a short completion.

124 models · rates checked . Price rank only, not a quality score.

Ranked by published rates

#ModelProviderContextInput / 1MOutput / 1MLong example10k requests
1Llama 3.2 1B (Groq)Groq128K$0.040/1M$0.040/1M$0.00012$1.2Cost
2Llama 3.2 3B (Groq)Groq128K$0.060/1M$0.060/1M$0.00018$1.8Cost
3Llama 3.2 3B (Together)Together AI128K$0.060/1M$0.060/1M$0.00018$1.8Cost
4Gemma 7B (Groq)Groq8K$0.070/1M$0.070/1M$0.00021$2.1Cost
5Llama 3.1 8B Instant (Groq)Groq128K$0.050/1M$0.080/1M$0.00021$2.1Cost
6Llama 3.2 3B (Fireworks)Fireworks128K$0.100/1M$0.100/1M$0.0003$3.00Cost
7Gemini 1.5 Flash-8BGoogle1M$0.037/1M$0.150/1M$0.000337$3.37Cost
8Command R7BCohere128K$0.037/1M$0.150/1M$0.000337$3.37Cost
9Ministral 8BMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
10Mistral NemoMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
11Pixtral 12BMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
12Mistral SmallMistral128K$0.060/1M$0.180/1M$0.00042$4.2Cost
13Llama 3.2 11B Vision (Groq)Groq128K$0.180/1M$0.180/1M$0.00054$5.4Cost
14Llama 3.1 8B (Together)Together AI128K$0.180/1M$0.180/1M$0.00054$5.4Cost
15Gemma 2 9B (Groq)Groq8K$0.200/1M$0.200/1M$0.0006$6.00Cost
16Llama Guard 3 8B (Groq)Groq8K$0.200/1M$0.200/1M$0.0006$6.00Cost
17Mistral 7B (Together)Together AI33K$0.200/1M$0.200/1M$0.0006$6.00Cost
18Llama 3.1 8B (Fireworks)Fireworks128K$0.200/1M$0.200/1M$0.0006$6.00Cost
19MythoMax L2 13B (Fireworks)Fireworks4K$0.200/1M$0.200/1M$0.0006$6.00Cost
20Mixtral 8x7B (Groq)Groq33K$0.240/1M$0.240/1M$0.00072$7.2Cost
21Mistral TinyMistral33K$0.250/1M$0.250/1M$0.00075$7.5Cost
22DeepSeek CoderDeepSeek128K$0.140/1M$0.280/1M$0.0007$7.00Cost
23DeepSeek V4 FlashDeepSeek1M$0.140/1M$0.280/1M$0.0007$7.00Cost
24DeepSeek VLDeepSeek4K$0.140/1M$0.280/1M$0.0007$7.00Cost
25Gemini 1.5 FlashGoogle1M$0.075/1M$0.300/1M$0.000675$6.75Cost
26Gemini 2.0 Flash-LiteGoogle1M$0.075/1M$0.300/1M$0.000675$6.75Cost
27Qwen2.5 7B (Together)Together AI33K$0.300/1M$0.300/1M$0.0009$9.00Cost
28Llama 4 Scout (Groq)Groq128K$0.110/1M$0.340/1M$0.00079$7.9Cost
29QwQ 32B (Groq)Groq128K$0.290/1M$0.390/1M$0.00107$10.7Cost
30GPT-5 NanoOpenAI128K$0.050/1M$0.400/1M$0.00085$8.5Cost
31GPT-4.1 NanoOpenAI1M$0.100/1M$0.400/1M$0.0009$9.00Cost
32Gemini 2.0 FlashGoogle1M$0.100/1M$0.400/1M$0.0009$9.00Cost
33Gemini 2.5 Flash-LiteGoogle1M$0.100/1M$0.400/1M$0.0009$9.00Cost
34Gemma 3 27BGoogle128K$0.200/1M$0.400/1M$0.001$10.00Cost
35DeepSeek ChatDeepSeek128K$0.280/1M$0.420/1M$0.00112$11.2Cost
36Grok 2 MinixAI131K$0.200/1M$0.500/1M$0.0012$12.00Cost
37Grok 4 FastxAI256K$0.200/1M$0.500/1M$0.0012$12.00Cost
38Grok 3 MinixAI131K$0.300/1M$0.500/1M$0.0013$13.00Cost
39Qwen3 32B (Groq)Groq131K$0.290/1M$0.590/1M$0.00147$14.7Cost
40GPT-4o miniOpenAI128K$0.150/1M$0.600/1M$0.00135$13.5Cost
41Command RCohere128K$0.150/1M$0.600/1M$0.00135$13.5Cost
42Command R (08-2024)Cohere128K$0.150/1M$0.600/1M$0.00135$13.5Cost
43Mistral SabaMistral33K$0.200/1M$0.600/1M$0.0014$14.00Cost
44DeepSeek R1 Distill Qwen 32BDeepSeek128K$0.300/1M$0.600/1M$0.0015$15.00Cost
45Command LightCohere4K$0.300/1M$0.600/1M$0.0015$15.00Cost
46Open Mixtral 8x7BMistral33K$0.700/1M$0.700/1M$0.0021$21.00Cost
47Llama 3.3 70B (Groq)Groq128K$0.590/1M$0.790/1M$0.00217$21.7Cost
48Qwen2.5 Coder 32B (Together)Together AI33K$0.800/1M$0.800/1M$0.0024$24.00Cost
49Llama 4 Maverick (Together)Together AI1M$0.270/1M$0.850/1M$0.00197$19.7Cost
50DeepSeek V4 ProDeepSeek1M$0.435/1M$0.870/1M$0.002175$21.75Cost
51Llama 4 Maverick (Fireworks)Fireworks1M$0.220/1M$0.880/1M$0.00198$19.8Cost
52Llama 3.1 70B (Together)Together AI128K$0.880/1M$0.880/1M$0.00264$26.4Cost
53Llama 3.3 70B (Together)Together AI128K$0.880/1M$0.880/1M$0.00264$26.4Cost
54DeepSeek R1 (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
55DeepSeek V3 (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
56Llama 3.1 70B (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
57Llama 3.3 70B (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
58Mixtral 8x22B (Fireworks)Fireworks66K$0.900/1M$0.900/1M$0.0027$27.00Cost
59Qwen2.5 32B (Fireworks)Fireworks33K$0.900/1M$0.900/1M$0.0027$27.00Cost
60Qwen2.5 72B (Fireworks)Fireworks33K$0.900/1M$0.900/1M$0.0027$27.00Cost
61Llama 3.3 70B SpecDec (Groq)Groq8K$0.590/1M$0.990/1M$0.00257$25.7Cost
62DeepSeek R1 Distill Llama 70B (Groq)Groq128K$0.750/1M$0.990/1M$0.00273$27.3Cost
63DeepSeek V3DeepSeek128K$0.270/1M$1.10/1M$0.00247$24.7Cost
64GPT-5.6 LunaOpenAI1.1M$0.200/1M$1.20/1M$0.0026$26.00Cost
65Mixtral 8x22B (Together)Together AI66K$1.20/1M$1.20/1M$0.0036$36.00Cost
66Qwen2.5 72B (Together)Together AI33K$1.20/1M$1.20/1M$0.0036$36.00Cost
67Claude 3 HaikuAnthropic200K$0.250/1M$1.25/1M$0.00275$27.5Cost
68DeepSeek V3 (Together)Together AI128K$1.25/1M$1.25/1M$0.00375$37.5Cost
69Gemini 3.1 Flash-LiteGoogle1M$0.250/1M$1.50/1M$0.00325$32.5Cost
70GPT-3.5 TurboOpenAI16K$0.500/1M$1.50/1M$0.0035$35.00Cost
71Gemini 1.0 ProGoogle33K$0.500/1M$1.50/1M$0.0035$35.00Cost
72Mistral Large 3Mistral128K$0.500/1M$1.50/1M$0.0035$35.00Cost
73GPT-4.1 MiniOpenAI1M$0.400/1M$1.60/1M$0.0036$36.00Cost
74GPT-5 MiniOpenAI128K$0.250/1M$2.00/1M$0.00425$42.5Cost
75Mistral Medium 3Mistral128K$0.400/1M$2.00/1M$0.0044$44.00Cost
76Command NightlyCohere128K$1.00/1M$2.00/1M$0.005$50.00Cost
77DeepSeek R1DeepSeek128K$0.550/1M$2.19/1M$0.00493$49.3Cost
78DeepSeek ReasonerDeepSeek128K$0.550/1M$2.19/1M$0.00493$49.3Cost
79Gemini 2.5 FlashGoogle1M$0.300/1M$2.50/1M$0.0053$53.00Cost
80Gemini 3 FlashGoogle1M$0.500/1M$3.00/1M$0.0065$65.00Cost
81CodestralMistral256K$1.00/1M$3.00/1M$0.007$70.00Cost
82Claude 3.5 HaikuAnthropic200K$0.800/1M$4.00/1M$0.0088$88.00Cost
83o1-miniOpenAI128K$1.10/1M$4.40/1M$0.0099$99.00Cost
84o3-miniOpenAI200K$1.10/1M$4.40/1M$0.0099$99.00Cost
85o4-miniOpenAI200K$1.10/1M$4.40/1M$0.0099$99.00Cost
86Claude Haiku 4.5Anthropic200K$1.00/1M$5.00/1M$0.011$110.00Cost
87Gemini 1.5 ProGoogle2M$1.25/1M$5.00/1M$0.01125$112.5Cost
88Gemini 2.0 Pro ExpGoogle2M$1.25/1M$5.00/1M$0.01125$112.5Cost
89Grok 4.6xAI500K$2.00/1M$6.00/1M$0.014$140.00Cost
90Mistral Large (2407)Mistral128K$2.00/1M$6.00/1M$0.014$140.00Cost
91Open Mixtral 8x22BMistral66K$2.00/1M$6.00/1M$0.014$140.00Cost
92DeepSeek R1 (Together)Together AI128K$3.00/1M$7.00/1M$0.017$170.00Cost
93GPT-4.1OpenAI1M$2.00/1M$8.00/1M$0.018$180.00Cost
94GPT-4.1 (2025-04-14)OpenAI1M$2.00/1M$8.00/1M$0.018$180.00Cost
95o3OpenAI200K$2.00/1M$8.00/1M$0.018$180.00Cost
96GPT-5OpenAI400K$1.25/1M$10.00/1M$0.02125$212.5Cost
97Gemini 2.5 ProGoogle1M$1.25/1M$10.00/1M$0.02125$212.5Cost
98Claude Sonnet 5Anthropic1M$2.00/1M$10.00/1M$0.022$220.00Cost
99Grok 2xAI131K$2.00/1M$10.00/1M$0.022$220.00Cost
100Grok 2 VisionxAI33K$2.00/1M$10.00/1M$0.022$220.00Cost
101GPT-4oOpenAI128K$2.50/1M$10.00/1M$0.0225$225.00Cost
102GPT-4o (2024-08-06)OpenAI128K$2.50/1M$10.00/1M$0.0225$225.00Cost
103Command ACohere256K$2.50/1M$10.00/1M$0.0225$225.00Cost
104Command R+Cohere128K$2.50/1M$10.00/1M$0.0225$225.00Cost
105Command R+ (08-2024)Cohere128K$2.50/1M$10.00/1M$0.0225$225.00Cost
106GPT-5.6 TerraOpenAI1.1M$2.00/1M$12.00/1M$0.026$260.00Cost
107Gemini 3.1 ProGoogle1M$2.00/1M$12.00/1M$0.026$260.00Cost
108Claude 3.5 SonnetAnthropic200K$3.00/1M$15.00/1M$0.033$330.00Cost
109Claude 3 SonnetAnthropic200K$3.00/1M$15.00/1M$0.033$330.00Cost
110Claude Sonnet 4.6Anthropic1M$3.00/1M$15.00/1M$0.033$330.00Cost
111Grok 3xAI131K$3.00/1M$15.00/1M$0.033$330.00Cost
112Grok 4xAI256K$3.00/1M$15.00/1M$0.033$330.00Cost
113ChatGPT-4o LatestOpenAI128K$5.00/1M$15.00/1M$0.035$350.00Cost
114Grok BetaxAI131K$5.00/1M$15.00/1M$0.035$350.00Cost
115Grok Vision BetaxAI8K$5.00/1M$15.00/1M$0.035$350.00Cost
116Claude 2.1Anthropic200K$8.00/1M$24.00/1M$0.056$560.00Cost
117Claude Opus 4.6Anthropic1M$5.00/1M$25.00/1M$0.055$550.00Cost
118Claude Opus 4.8Anthropic1M$5.00/1M$25.00/1M$0.055$550.00Cost
119GPT-5.6 SolOpenAI1.1M$5.00/1M$30.00/1M$0.065$650.00Cost
120GPT-4 TurboOpenAI128K$10.00/1M$30.00/1M$0.07$700.00Cost
121o1OpenAI200K$15.00/1M$60.00/1M$0.135$1350.00Cost
122o1-previewOpenAI128K$15.00/1M$60.00/1M$0.135$1350.00Cost
123GPT-4OpenAI8K$30.00/1M$60.00/1M$0.15$1500.00Cost
124Claude 3 OpusAnthropic200K$15.00/1M$75.00/1M$0.165$1650.00Cost

Related tools

LLM cost calculator · Cheapest LLM API · Compare model prices · Cache savings · Batch pricing

Related guides

How LLM API pricing works · How prompt cost is calculated · Prompt caching explained · Exact vs Approx

Sources and references

Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.

Pricing rank FAQ

Why rank by output?
Long answers are billed on output tokens, which are usually more expensive than input. A cheap input rate can still be a costly chat model.
What is the long example?
1000 input tokens and 2000 output tokens at standard rates, no cache and no batch.
Can I change the mix?
Yes. Open Cost on any row and set your own output size and monthly volume.