Free AI API cost calculator. Compare real per-request costs across 24 models from OpenAI, Anthropic, Google, DeepSeek, Mistral, xAI, Meta, and Amazon, with the caching, batch, and reasoning-token math most calculators skip.
| Model | $ / 1M in · out | Cost / request | Cost / 1K requests | Cost / month |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | $5.00 · $30.00 | $0.017750 | $17.750 | $1775.00 |
| GPT-5.6 Terra OpenAI | $2.50 · $15.00 | $0.008875 | $8.875 | $887.50 |
| GPT-5.6 Luna OpenAI | $1.00 · $6.00 | $0.003550 | $3.550 | $355.00 |
| GPT-5.4 mini OpenAI | $0.75 · $4.50 | $0.002662 | $2.663 | $266.25 |
| GPT-5.4 nano OpenAI | $0.20 · $1.25 | $0.000735 | $0.735 | $73.50 |
| Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3× | $10.00 · $50.00 | $0.031125 | $31.125 | $3112.50 |
| Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3× | $5.00 · $25.00 | $0.015562 | $15.562 | $1556.25 |
| Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled | $2.00 · $10.00 | $0.009338 | $9.338 | $933.75 |
| Claude Haiku 4.5 Anthropic (Claude) | $1.00 · $5.00 | $0.003112 | $3.112 | $311.25 |
| Gemini 3.1 Pro (Preview) Google (Gemini API) | $2.00 · $12.00 | $0.007100 | $7.100 | $710.00 |
| Gemini 3.5 Flash Google (Gemini API) | $1.50 · $9.00 | $0.005325 | $5.325 | $532.50 |
| Gemini 2.5 Flash-Lite Google (Gemini API) | $0.10 · $0.40 | $0.000255 | $0.255 | $25.50 |
| DeepSeek V4 Pro DeepSeek | $0.43 · $0.87 | $0.000654 | $0.654 | $65.43 |
| DeepSeek V4 Flash DeepSeek | $0.14 · $0.28 | $0.000211 | $0.211 | $21.14 |
| Mistral Medium 3.5 Mistral | $1.50 · $7.50 | $0.005250 | $5.250 | $525.00 |
| Mistral Large 3 Mistral | $0.50 · $1.50 | $0.001025 | $1.025 | $102.50 |
| Mistral Small 4 Mistral | $0.15 · $0.60 | $0.000382 | $0.382 | $38.25 |
| Grok 4.5 xAI (Grok) | $2.00 · $6.00 | $0.004250 | $4.250 | $425.00 |
| Grok 4.3 xAI (Grok) | $1.25 · $2.50 | $0.001975 | $1.975 | $197.50 |
| Llama 4 Scout Meta Llama (hosted) | $0.10 · $0.30 | $0.000250 | $0.250 | $25.00 |
| Llama 4 Maverick Meta Llama (hosted) | $0.15 · $0.60 | $0.000450 | $0.450 | $45.00 |
| Amazon Nova Pro Amazon (Nova on Bedrock) | $0.80 · $3.20 | $0.002400 | $2.400 | $240.00 |
| Amazon Nova Lite Amazon (Nova on Bedrock) | $0.06 · $0.24 | $0.000180 | $0.180 | $18.00 |
| Amazon Nova Micro Amazon (Nova on Bedrock) | $0.04 · $0.14 | $0.000105 | $0.105 | $10.50 |
The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.
The two numbers on a pricing page, dollars per million input and output tokens, hide most of what determines a real AI bill. Output tokens cost 3–6x input tokens everywhere, so the balance of input to output tokens matters more than either list price. Prompt caching changes input economics by up to 99% on some providers, and the three caching systems (implicit discounts, paid cache writes, storage-billed caches) aren’t comparable by glancing at a rate card. Batch APIs halve costs on six of the eight providers here for anything asynchronous. Reasoning models bill their invisible thinking tokens as output. And prices themselves move constantly. This year alone OpenAI shipped three flagship generations at different prices, Anthropic launched Claude Sonnet 5 with introductory pricing that rises on September 1, and Google’s Gemini 3.1 Pro charges double above 200K input tokens.
The calculator above models all of it. Pick a workload preset or enter your own token counts, and the table recomputes what a request actually costs on all 24 models, with the cheapest one for your inputs highlighted.
A chat assistant averages about 1,000 input tokens per request (half of it a cacheable system prompt) against 500 output. RAG flips the ratio, as retrieved chunks push input to 6,000+ tokens per query while answers stay short. Agent workflows are the expensive surprise. A single task can burn 150,000 cumulative input tokens across 20–50 model calls, which is why cache rates and reasoning multipliers dominate agent economics, and why one intensive session on a flagship model costs $2–5. Summarization is input-heavy and asynchronous, making it the ideal candidate for batch pricing.
Production guides consistently recommend multiplying naive estimates by 1.5–2x for retries, tool-call overhead, and growth in context length. Treat the calculator’s output as a floor, not a forecast.
Base per-token prices come from a live feed and refresh weekly, the same pipeline approach used by the freshest tools in this space, because a stale calculator is worse than none. Everything the feed can’t express is curated by hand against official API pricing documentation (never marketing pages, which go stale). That covers caching mechanics per provider, batch discounts, Gemini’s long-context tier, DeepSeek’s first-party prices where routed rates differ, and dated future changes like the Sonnet 5 increase, which this calculator will apply automatically on the day it takes effect.
Where a provider doesn’t publish something, we show a dash rather than assuming zero. And when you’re done estimating costs, the harder question of what to charge has its own tool. The AI margin calculator linked below inverts this math into a price per seat.
Base per-token prices refresh weekly from the OpenRouter model feed (last refresh 2026-07-11). Everything the feed can’t see, such as caching mechanics, batch discounts, long-context tiers, and announced price changes, is verified by hand against each provider’s official API pricing page.