Free developer calculator
LLM Cost Calculator
Enter your real token workload to estimate cost per request, per day, per month, and per 1,000 requests.
How LLM API cost is calculated
Standard input tokens use the model's input rate, eligible cache hits use its cached-input rate, and generated tokens use the output rate. Output often costs more because generation requires sequential inference, so a model's input unit price alone does not describe a complete workload.
Prices come from the same local model dataset used by the pricing pages. The calculator applies the selected model's verified rates, token mix, request volume, and cache usage rather than comparing unit price alone.
Worked monthly example
For Gemini 3.6 Flash, 100,000 monthly requests with 1,000 input tokens and 300 output tokens per request produce this estimate from the shared pricing calculation.
- Monthly input tokens
- 100,000,000
- Monthly output tokens
- 30,000,000
- Estimated monthly API cost
- $187.50
1,000 monthly requests
$3.04
100,000 monthly requests
$303.75
1,000,000 monthly requests
$3,037.50
Scale cards use the same 2,000 input / 500 cached / 500 output token request profile.
Practical ways to reduce cost
- Trim repeated instructions and retrieved context.
- Cache stable prompt prefixes when the provider supports it.
- Set realistic output limits and route simpler work to cheaper models.
- Compare total workload cost, quality, latency, and context needs together.
Common questions
What counts as input?
Your prompt, system instructions, conversation history, retrieved documents, and tool results all contribute input tokens.
Are cached tokens free?
No. Providers may bill eligible cache hits at a lower rate, which the calculator applies when the selected model lists one.
Is this the same as a subscription price?
No. This calculator estimates usage-based API charges, not consumer chat or coding-tool subscriptions.
Why can actual invoices differ?
Provider tokenization, pricing tiers, batch modes, cache eligibility, and billing updates can change the final charge.
Estimate a prompt with the Token Calculator, check fit with the Context Window Calculator, compare all model API prices, or review Claude Sonnet 5 pricing.
For GPT-5.6 models, V1 quotes Standard short-context pricing through 272,000 input tokens per request. Larger requests use OpenAI long-context pricing and are not quoted at the short-context rate.