Home / Tools / Context window
Context window calculator Budget system prompt, history, retrieved docs, the user message, and reserved output against a model window. See whether the request fits before you call the API.
Tokenizer · Cost calculator · Long context ranks · Context windows guide
Token budget Paste document
Model OpenAI. ChatGPT-4o Latest (128,000 tok) OpenAI. GPT-3.5 Turbo (16,384 tok) OpenAI. GPT-4 (8,192 tok) OpenAI. GPT-4.1 (1,000,000 tok) OpenAI. GPT-4.1 (2025-04-14) (1,000,000 tok) OpenAI. GPT-4.1 Mini (1,000,000 tok) OpenAI. GPT-4.1 Nano (1,000,000 tok) OpenAI. GPT-4 Turbo (128,000 tok) OpenAI. GPT-4o (128,000 tok) OpenAI. GPT-4o (2024-08-06) (128,000 tok) OpenAI. GPT-4o mini (128,000 tok) OpenAI. GPT-5 (400,000 tok) OpenAI. GPT-5 Mini (128,000 tok) OpenAI. GPT-5 Nano (128,000 tok) OpenAI. GPT-5.6 Sol (1,050,000 tok) OpenAI. GPT-5.6 Terra (1,050,000 tok) OpenAI. GPT-5.6 Luna (1,050,000 tok) OpenAI. o1 (200,000 tok) OpenAI. o1-mini (128,000 tok) OpenAI. o1-preview (128,000 tok) OpenAI. o3 (200,000 tok) OpenAI. o3-mini (200,000 tok) OpenAI. o4-mini (200,000 tok) Anthropic. Claude 2.1 (200,000 tok) Anthropic. Claude 3.5 Haiku (200,000 tok) Anthropic. Claude 3.5 Sonnet (200,000 tok) Anthropic. Claude 3 Haiku (200,000 tok) Anthropic. Claude 3 Opus (200,000 tok) Anthropic. Claude 3 Sonnet (200,000 tok) Anthropic. Claude Haiku 4.5 (200,000 tok) Anthropic. Claude Opus 4.6 (1,000,000 tok) Anthropic. Claude Opus 4.8 (1,000,000 tok) Anthropic. Claude Sonnet 4.6 (1,000,000 tok) Anthropic. Claude Sonnet 5 (1,000,000 tok) Google. Gemini 1.0 Pro (32,768 tok) Google. Gemini 1.5 Flash (1,000,000 tok) Google. Gemini 1.5 Flash-8B (1,000,000 tok) Google. Gemini 1.5 Pro (2,000,000 tok) Google. Gemini 2.0 Flash (1,000,000 tok) Google. Gemini 2.0 Flash-Lite (1,000,000 tok) Google. Gemini 2.0 Pro Exp (2,000,000 tok) Google. Gemini 2.5 Flash (1,000,000 tok) Google. Gemini 2.5 Flash-Lite (1,000,000 tok) Google. Gemini 2.5 Pro (1,000,000 tok) Google. Gemini 3.1 Flash-Lite (1,000,000 tok) Google. Gemini 3.1 Pro (1,000,000 tok) Google. Gemini 3 Flash (1,000,000 tok) Google. Gemma 3 27B (128,000 tok) DeepSeek. DeepSeek Chat (128,000 tok) DeepSeek. DeepSeek Coder (128,000 tok) DeepSeek. DeepSeek R1 (128,000 tok) DeepSeek. DeepSeek R1 Distill Qwen 32B (128,000 tok) DeepSeek. DeepSeek Reasoner (128,000 tok) DeepSeek. DeepSeek V3 (128,000 tok) DeepSeek. DeepSeek V4 Flash (1,000,000 tok) DeepSeek. DeepSeek V4 Pro (1,000,000 tok) DeepSeek. DeepSeek VL (4,096 tok) xAI. Grok 2 (131,072 tok) xAI. Grok 2 Mini (131,072 tok) xAI. Grok 2 Vision (32,768 tok) xAI. Grok 3 (131,072 tok) xAI. Grok 3 Mini (131,072 tok) xAI. Grok 4.6 (500,000 tok) xAI. Grok 4 (256,000 tok) xAI. Grok 4 Fast (256,000 tok) xAI. Grok Beta (131,072 tok) xAI. Grok Vision Beta (8,192 tok) Mistral. Codestral (256,000 tok) Mistral. Ministral 8B (128,000 tok) Mistral. Mistral Large (2407) (128,000 tok) Mistral. Mistral Large 3 (128,000 tok) Mistral. Mistral Medium 3 (128,000 tok) Mistral. Mistral Nemo (128,000 tok) Mistral. Mistral Saba (32,768 tok) Mistral. Mistral Small (128,000 tok) Mistral. Mistral Tiny (32,768 tok) Mistral. Open Mixtral 8x22B (65,536 tok) Mistral. Open Mixtral 8x7B (32,768 tok) Mistral. Pixtral 12B (128,000 tok) Groq. DeepSeek R1 Distill Llama 70B (Groq) (128,000 tok) Groq. Gemma 7B (Groq) (8,192 tok) Groq. Gemma 2 9B (Groq) (8,192 tok) Groq. Llama 3.1 8B Instant (Groq) (128,000 tok) Groq. Llama 3.2 11B Vision (Groq) (128,000 tok) Groq. Llama 3.2 1B (Groq) (128,000 tok) Groq. Llama 3.2 3B (Groq) (128,000 tok) Groq. Llama 3.3 70B (Groq) (128,000 tok) Groq. Llama 3.3 70B SpecDec (Groq) (8,192 tok) Groq. Llama 4 Scout (Groq) (128,000 tok) Groq. Llama Guard 3 8B (Groq) (8,192 tok) Groq. Mixtral 8x7B (Groq) (32,768 tok) Groq. QwQ 32B (Groq) (128,000 tok) Groq. Qwen3 32B (Groq) (131,072 tok) Cohere. Command A (256,000 tok) Cohere. Command Light (4,096 tok) Cohere. Command Nightly (128,000 tok) Cohere. Command R (128,000 tok) Cohere. Command R (08-2024) (128,000 tok) Cohere. Command R+ (128,000 tok) Cohere. Command R+ (08-2024) (128,000 tok) Cohere. Command R7B (128,000 tok) Together AI. DeepSeek R1 (Together) (128,000 tok) Together AI. DeepSeek V3 (Together) (128,000 tok) Together AI. Llama 3.1 70B (Together) (128,000 tok) Together AI. Llama 3.1 8B (Together) (128,000 tok) Together AI. Llama 3.2 3B (Together) (128,000 tok) Together AI. Llama 3.3 70B (Together) (128,000 tok) Together AI. Llama 4 Maverick (Together) (1,000,000 tok) Together AI. Mistral 7B (Together) (32,768 tok) Together AI. Mixtral 8x22B (Together) (65,536 tok) Together AI. Qwen2.5 72B (Together) (32,768 tok) Together AI. Qwen2.5 7B (Together) (32,768 tok) Together AI. Qwen2.5 Coder 32B (Together) (32,768 tok) Fireworks. DeepSeek R1 (Fireworks) (128,000 tok) Fireworks. DeepSeek V3 (Fireworks) (128,000 tok) Fireworks. Llama 3.1 70B (Fireworks) (128,000 tok) Fireworks. Llama 3.1 8B (Fireworks) (128,000 tok) Fireworks. Llama 3.2 3B (Fireworks) (128,000 tok) Fireworks. Llama 3.3 70B (Fireworks) (128,000 tok) Fireworks. Llama 4 Maverick (Fireworks) (1,000,000 tok) Fireworks. Mixtral 8x22B (Fireworks) (65,536 tok) Fireworks. MythoMax L2 13B (Fireworks) (4,096 tok) Fireworks. Qwen2.5 32B (Fireworks) (32,768 tok) Fireworks. Qwen2.5 72B (Fireworks) (32,768 tok) OpenAI · window 1,000,000 tokens · Exact counts · rates checked 2026-08-12
Fit Fits with room
Input total 9,100
With reserved output 9,612
Remaining 890,388 (1% used) Models that fit this budget Same total (9,612 tokens) against published windows. Smallest fitting windows first.
Cheapest long context · Cost for GPT-4.1 · Tokenizer
How to use this Pick the model you plan to call. Enter token budgets for system, history, docs, and the user message, or paste a document to count the docs bucket. Reserve output tokens for the reply you actually need. Read the fit verdict, then open the cost calculator if the request fits. Sources and references Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.
FAQ
What is a context window? The maximum tokens a model can hold in one request. Input and the reserved reply share the same budget.
Why reserve output tokens? The reply consumes the same window as your prompt. If you fill the window with input, the answer gets truncated or the API rejects the call.
Are paste counts Exact? OpenAI encodings are Exact in the browser. Other providers are Approx. Confirm critical budgets with the provider count API before production.
What does the safety margin do? It shrinks the usable window for planning (default 10%). That buffer covers tool wrappers, formatting tokens, and tokenizer drift on Approx models.