Home / Tools / GPU LLM cost

GPU / LLM cost calculator

Compare catalog LLM API spend to self-host or cloud GPU cost at your volume. See break-even requests, capacity headroom, and which side wins this month.

API side (TokenCalculator catalog)

OpenAI · input $0.400/1M · output $1.60/1M · rates checked 2026-08-12

250,000,000 tokens / month (2,500 per request).

GPU side (editable reference defaults)

Reference defaults. Spot and committed rates differ a lot.

Include monitoring, patches, and on-call. Idle hours still burn rental and much of ops.

API / request
$0.0016
API / month
$160.00
GPU stack / month
$1482.96
Delta (API minus GPU)
$-1322.96 (API cheaper)
Break-even requests / month
926,850
Break-even tokens / month
2,317,125,000
GPU capacity (tokens / month)
2,851,200,000
Capacity vs load
9% (OK headroom)
LineMonthly
API (catalog rates)$160.00
GPU rental$720.00
Electricity$12.96
Ops / MLOps$750.00

How to use this

  1. Pick the API model you would pay today (or a fair open-model host comparison).
  2. Enter tokens per request and monthly request volume.
  3. Choose a GPU preset, edit $/hr or purchase terms, hours/day, tok/s, power, and ops.
  4. Read who wins this month, the break-even volume, and whether capacity holds.

Open GPT-4.1 Mini in cost calculator · Self-host vs API · Break-even guide · When to self-host

Related tools

API cost calculator · Cheapest API models

Related guides

GPU LLM cost calculator: compare API spend to self-host GPUs · Self-host vs API · GPU break-even · When to self-host · GPU utilization

Sources and references

Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.

FAQ

Is it cheaper to self-host an LLM or use an API?
It depends on sustained volume, GPU utilization, ops cost, and which API you replace. Low volume usually favors APIs. High steady volume can favor a busy GPU. Run your numbers here.
At what volume does self-hosting beat API pricing?
Break-even is roughly monthly GPU cost ÷ API cost per request (or per token mix). The calculator prints that volume for your inputs.
Are GPU $/hr and tokens/sec from the TokenCalculator catalog?
No. Those are editable reference defaults. Always verify with your cloud vendor and a measured throughput on your model stack.
Should I compare a self-hosted Llama to Claude or GPT frontier APIs?
Only for budget brainstorming. Quality, tools, and latency differ. Prefer comparing to open-model hosts (or the same open weights via API) when quality must match.