AI API Cost Calculator With Live LLM Prices

Free AI API cost calculator. Compare real per-request costs across 24 models from OpenAI, Anthropic, Google, DeepSeek, Mistral, xAI, Meta, and Amazon, with the caching, batch, and reasoning-token math most calculators skip.

Model$ / 1M in · outCost / requestCost / 1K requestsCost / month
GPT-5.6 Sol OpenAI$5.00 · $30.00$0.017750$17.750$1775.00
GPT-5.6 Terra OpenAI$2.50 · $15.00$0.008875$8.875$887.50
GPT-5.6 Luna OpenAI$1.00 · $6.00$0.003550$3.550$355.00
GPT-5.4 mini OpenAI$0.75 · $4.50$0.002662$2.663$266.25
GPT-5.4 nano OpenAI$0.20 · $1.25$0.000735$0.735$73.50
Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3×$10.00 · $50.00$0.031125$31.125$3112.50
Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3×$5.00 · $25.00$0.015562$15.562$1556.25
Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled$2.00 · $10.00$0.009338$9.338$933.75
Claude Haiku 4.5 Anthropic (Claude)$1.00 · $5.00$0.003112$3.112$311.25
Gemini 3.1 Pro (Preview) Google (Gemini API)$2.00 · $12.00$0.007100$7.100$710.00
Gemini 3.5 Flash Google (Gemini API)$1.50 · $9.00$0.005325$5.325$532.50
Gemini 2.5 Flash-Lite Google (Gemini API)$0.10 · $0.40$0.000255$0.255$25.50
DeepSeek V4 Pro DeepSeek$0.43 · $0.87$0.000654$0.654$65.43
DeepSeek V4 Flash DeepSeek$0.14 · $0.28$0.000211$0.211$21.14
Mistral Medium 3.5 Mistral$1.50 · $7.50$0.005250$5.250$525.00
Mistral Large 3 Mistral$0.50 · $1.50$0.001025$1.025$102.50
Mistral Small 4 Mistral$0.15 · $0.60$0.000382$0.382$38.25
Grok 4.5 xAI (Grok)$2.00 · $6.00$0.004250$4.250$425.00
Grok 4.3 xAI (Grok)$1.25 · $2.50$0.001975$1.975$197.50
Llama 4 Scout Meta Llama (hosted)$0.10 · $0.30$0.000250$0.250$25.00
Llama 4 Maverick Meta Llama (hosted)$0.15 · $0.60$0.000450$0.450$45.00
Amazon Nova Pro Amazon (Nova on Bedrock)$0.80 · $3.20$0.002400$2.400$240.00
Amazon Nova Lite Amazon (Nova on Bedrock)$0.06 · $0.24$0.000180$0.180$18.00
Amazon Nova Micro Amazon (Nova on Bedrock)$0.04 · $0.14$0.000105$0.105$10.50

The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.

Why list prices mislead

The two numbers on a pricing page, dollars per million input and output tokens, hide most of what determines a real AI bill. Output tokens cost 3–6x input tokens everywhere, so the balance of input to output tokens matters more than either list price. Prompt caching changes input economics by up to 99% on some providers, and the three caching systems (implicit discounts, paid cache writes, storage-billed caches) aren’t comparable by glancing at a rate card. Batch APIs halve costs on six of the eight providers here for anything asynchronous. Reasoning models bill their invisible thinking tokens as output. And prices themselves move constantly. This year alone OpenAI shipped three flagship generations at different prices, Anthropic launched Claude Sonnet 5 with introductory pricing that rises on September 1, and Google’s Gemini 3.1 Pro charges double above 200K input tokens.

The calculator above models all of it. Pick a workload preset or enter your own token counts, and the table recomputes what a request actually costs on all 24 models, with the cheapest one for your inputs highlighted.

How the presets map to real workloads

A chat assistant averages about 1,000 input tokens per request (half of it a cacheable system prompt) against 500 output. RAG flips the ratio, as retrieved chunks push input to 6,000+ tokens per query while answers stay short. Agent workflows are the expensive surprise. A single task can burn 150,000 cumulative input tokens across 20–50 model calls, which is why cache rates and reasoning multipliers dominate agent economics, and why one intensive session on a flagship model costs $2–5. Summarization is input-heavy and asynchronous, making it the ideal candidate for batch pricing.

Production guides consistently recommend multiplying naive estimates by 1.5–2x for retries, tool-call overhead, and growth in context length. Treat the calculator’s output as a floor, not a forecast.

How we keep this honest

Base per-token prices come from a live feed and refresh weekly, the same pipeline approach used by the freshest tools in this space, because a stale calculator is worse than none. Everything the feed can’t express is curated by hand against official API pricing documentation (never marketing pages, which go stale). That covers caching mechanics per provider, batch discounts, Gemini’s long-context tier, DeepSeek’s first-party prices where routed rates differ, and dated future changes like the Sonnet 5 increase, which this calculator will apply automatically on the day it takes effect.

Where a provider doesn’t publish something, we show a dash rather than assuming zero. And when you’re done estimating costs, the harder question of what to charge has its own tool. The AI margin calculator linked below inverts this math into a price per seat.

Price sources & freshness

Base per-token prices refresh weekly from the OpenRouter model feed (last refresh 2026-07-11). Everything the feed can’t see, such as caching mechanics, batch discounts, long-context tiers, and announced price changes, is verified by hand against each provider’s official API pricing page.

  • OpenAI official pricing, details last verified 2026-07-11 (The GPT-5.6 family (Sol/Terra/Luna) launched Jul 9, 2026. Its three-tier naming replaces the mini/nano ladder at the top end.)
  • Anthropic (Claude) official pricing, details last verified 2026-07-11 (Opus 4.7+, Sonnet 5, and Fable 5 use a new tokenizer producing up to roughly 30% more tokens for the same text, so cross-provider per-token comparisons understate Claude cost slightly.)
  • Google (Gemini API) official pricing, details last verified 2026-07-11 (Gemini 3.1 Pro doubles input pricing above 200K input tokens, the only two-tier context pricing in this comparison.)
  • DeepSeek official pricing, details last verified 2026-07-11 (Prices below are DeepSeek's first-party API. OpenRouter routes some DeepSeek models through discounted hosts at different prices.)
  • Mistral official pricing, details last verified 2026-07-11 (Verify against the API pricing page, since Mistral's marketing pages have shown stale prices for retired models.)
  • xAI (Grok) official pricing, details last verified 2026-07-11
  • Meta Llama (hosted) official pricing, details last verified 2026-07-11 (These are open-weight models, so the same model costs different amounts on different hosts such as Groq, DeepInfra, and Bedrock. Prices shown are OpenRouter's routed rate, indicative rather than contractual.)
  • Amazon (Nova on Bedrock) official pricing, details last verified 2026-07-11 (Bedrock also hosts Anthropic, Mistral, and DeepSeek models, with Claude at first-party prices and DeepSeek at different ones. Only Amazon's own Nova line is listed here.)

Frequently asked questions

Most calculators multiply list prices by token counts and stop. This one models what actually determines your bill, including prompt caching (three different mechanics across providers), batch discounts, long-context tier pricing, reasoning tokens billed as output, and announced future price changes. Base prices refresh weekly from a live feed, and the hand-verified details link to official pricing pages with verification dates.

In production, 30–90% of input tokens are repeated system prompts, tool definitions, and accumulated context. Cached input costs roughly 10% of the base rate on OpenAI and Anthropic and as little as 1–2% on DeepSeek, so a calculator that ignores caching can overstate input costs by 5–10x for agent workloads.

Reasoning models think before answering, and every provider bills those invisible thinking tokens at the output rate. Depending on the task, that multiplies effective output cost by 1.5–5x. The calculator applies a reasoning multiplier only to models that bill this way.

For a typical chat workload, budget models like Gemini 2.5 Flash-Lite, Amazon Nova Micro, and Llama 4 Scout compute cheapest per request, but the answer changes with your token mix, cache rate, and context size. Gemini 3.1 Pro doubles its input price above 200K input tokens, and DeepSeek’s cache pricing rewards repetitive workloads. Enter your own numbers above rather than trusting any static list.

Base per-token prices refresh weekly from the OpenRouter model feed, and the page shows the last refresh date. Structural details the feed can’t see, such as caching mechanics, batch discounts, and scheduled changes like Claude Sonnet 5’s September 1 price increase, are verified by hand against official pricing pages, each with its own verification date.

No, v1 covers text-token API pricing only. Multimodal rates (image and audio tokens), fine-tuning, priority tiers, and provisioned-throughput pricing follow different schedules. Where they materially change a provider’s economics we note it in that provider’s page.

Work backwards from a target gross margin. Healthy AI products keep API costs to 15–30% of revenue, which means charging roughly 3.5–5x your token cost. Our AI margin calculator does that inversion for you. Pick a model, enter usage per user, and it computes the price per seat that hits 70% or 80% margin.