OpenAI API Cost Calculator for GPT-5.6 Pricing (2026)

Calculate OpenAI API costs for GPT-5.6 Sol, Terra, and Luna with live per-request pricing, cache and batch discounts, and worked examples per workload.

Model$ / 1M in · outCost / requestCost / 1K requestsCost / month
GPT-5.6 Sol OpenAI$5.00 · $30.00$0.017750$17.750$1775.00
GPT-5.6 Terra OpenAI$2.50 · $15.00$0.008875$8.875$887.50
GPT-5.6 Luna OpenAI$1.00 · $6.00$0.003550$3.550$355.00
GPT-5.4 mini OpenAI$0.75 · $4.50$0.002662$2.663$266.25
GPT-5.4 nano OpenAI$0.20 · $1.25$0.000735$0.735$73.50
Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3×$10.00 · $50.00$0.031125$31.125$3112.50
Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3×$5.00 · $25.00$0.015562$15.562$1556.25
Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled$2.00 · $10.00$0.009338$9.338$933.75
Claude Haiku 4.5 Anthropic (Claude)$1.00 · $5.00$0.003112$3.112$311.25
Gemini 3.1 Pro (Preview) Google (Gemini API)$2.00 · $12.00$0.007100$7.100$710.00
Gemini 3.5 Flash Google (Gemini API)$1.50 · $9.00$0.005325$5.325$532.50
Gemini 2.5 Flash-Lite Google (Gemini API)$0.10 · $0.40$0.000255$0.255$25.50
DeepSeek V4 Pro DeepSeek$0.43 · $0.87$0.000654$0.654$65.43
DeepSeek V4 Flash DeepSeek$0.14 · $0.28$0.000211$0.211$21.14
Mistral Medium 3.5 Mistral$1.50 · $7.50$0.005250$5.250$525.00
Mistral Large 3 Mistral$0.50 · $1.50$0.001025$1.025$102.50
Mistral Small 4 Mistral$0.15 · $0.60$0.000382$0.382$38.25
Grok 4.5 xAI (Grok)$2.00 · $6.00$0.004250$4.250$425.00
Grok 4.3 xAI (Grok)$1.25 · $2.50$0.001975$1.975$197.50
Llama 4 Scout Meta Llama (hosted)$0.10 · $0.30$0.000250$0.250$25.00
Llama 4 Maverick Meta Llama (hosted)$0.15 · $0.60$0.000450$0.450$45.00
Amazon Nova Pro Amazon (Nova on Bedrock)$0.80 · $3.20$0.002400$2.400$240.00
Amazon Nova Lite Amazon (Nova on Bedrock)$0.06 · $0.24$0.000180$0.180$18.00
Amazon Nova Micro Amazon (Nova on Bedrock)$0.04 · $0.14$0.000105$0.105$10.50

The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.

How OpenAI pricing works

  • Prompt caching: Cached input billed at about 10% of base with no explicit write step
  • Batch API: −50% on asynchronous workloads.
  • The GPT-5.6 family (Sol/Terra/Luna) launched Jul 9, 2026. Its three-tier naming replaces the mini/nano ladder at the top end.
  • Regional data-residency processing adds ~10% (not modeled).
  • Source: official pricing page, details last verified 2026-07-11.

What the GPT-5.6 ladder costs per request

OpenAI’s GPT-5.6 family, launched July 9, 2026, prices its three general-purpose tiers in fixed ratios, with Sol at $5 input / $30 output per million tokens, Terra at exactly half ($2.50/$15), and Luna at one-fifth ($1/$6). Because the ratios are constant, per-request costs scale predictably across the three workload profiles the calculator models.

  • A chat exchange (1,000 input / 500 output tokens, 50% cache hits) runs $0.01775 on Sol, $0.008875 on Terra, and $0.00355 on Luna.
  • A RAG query (6,000 input / 400 output, 30% cached) costs $0.0339 on Sol, $0.01695 on Terra, and $0.00678 on Luna.
  • An agent task (150K input / 8K output, 80% cached, reasoning-heavy) comes to $0.63 on Sol, $0.315 on Terra, and $0.126 on Luna.

At 100,000 chat requests a month, that is $1,775 on Sol, $887.50 on Terra, or $355 on Luna. GPT-5.4 mini ($0.75/$4.50) still undercuts the whole 5.6 ladder at $0.002662 per chat exchange, exactly 25% below Luna, and cheaper rows exist outside OpenAI entirely, such as Grok 4.5 at $0.00425 and DeepSeek V4 Pro at $0.000654 for the same exchange.

When Terra’s 2.5x premium over Luna pays off

Terra costs exactly 2.5x Luna on every line item, so the question is whether accuracy gains justify a 2.5x bill. At 1,000 agent tasks per month the gap is $189 ($315 versus $126), and at 10,000 RAG queries it is $101.70 ($169.50 versus $67.80). Run evals both ways. If Luna passes, the premium buys nothing, and if Terra prevents even a small share of failed agent runs that would be retried at full cost, it can pay for itself. Sol’s further 2x step over Terra follows the same logic at higher stakes.

Reasoning tokens are the output multiplier you don’t see

All five GPT-5.6 models are reasoning models, and thinking tokens bill at the output rate without appearing in the response. The agent-task figures above assume a 1.75x reasoning multiplier, so 8,000 visible output tokens bill as 14,000. On Sol, that output line is $0.42 of the $0.63 total, two-thirds of the cost of a task whose input is nearly 19x larger than its visible output. Because output costs 6x input across the entire ladder, tuning reasoning effort down on tasks that don’t need deep deliberation usually saves more than trimming prompts.

The 5.4 to 5.5 to 5.6 price whiplash

OpenAI changed its price structure twice in three months this year. GPT-5.5, released in April 2026, doubled flagship pricing to $5/$30 and shipped with no mid-tier, leaving GPT-5.4 mini ($0.75/$4.50) and nano ($0.20/$1.25) as the only budget options. The GPT-5.6 launch on July 9 kept the $5/$30 flagship rate for Sol but restored the middle of the ladder with Terra and Luna. Teams that hardcoded a flagship model ID in April absorbed a 2x cost increase overnight, while teams that route the model name through configuration re-ran their evals and switched. The practical lesson is to treat the model ID as a config value, re-benchmark on every release, and let current prices drive the decision rather than a constant committed to your codebase six months ago.

Stacking cache and batch for async workloads

The two discounts multiply. Cached input bills at roughly 10% of base, and the Batch API halves everything, so a cached input token inside a batch job costs about 5% of list price. Work it through the agent task on Sol. At list price with no caching, 150K input tokens cost $0.75 and 14K billed output tokens cost $0.42, or $1.17 per task. An 80% cache-hit rate cuts that to $0.63, and running it through the Batch API cuts it to $0.315, a 73% total reduction. Push cache hits above 90% and the combined saving exceeds 75%. The constraints are concrete. Batch jobs complete within a 24-hour window rather than in real time, and caching only hits when prompt prefixes are byte-stable across requests, so put system prompts and tool definitions first and variable content last.

Picking a model by workload

For high-volume simple chat, GPT-5.4 mini at $0.002662 per exchange is the cheapest OpenAI option, with Luna the step up when reasoning or long context matters. For RAG, Luna’s $0.00678 per query covers most retrieval-grounded answering, and Terra’s $0.01695 is the upgrade when synthesis across many retrieved chunks needs to be reliable. For agents, a common pattern is Luna for routine steps with escalation to Sol for planning. At $0.126 versus $0.63 per task, an 80/20 split averages $0.227 per task, roughly a third of running everything on Sol. All five 5.6 models share context windows of about 1M tokens at a flat rate, so context length never forces a price-tier change the way it does on Gemini 3.1 Pro’s two-tier pricing.

Frequently asked questions

A typical chat exchange (1,000 input tokens, 500 output tokens, 50% cache hits) costs $0.01775 on GPT-5.6 Sol, $0.008875 on Terra, and $0.00355 on Luna. GPT-5.4 mini handles the same exchange for $0.002662. At 100,000 requests a month, that spans roughly $266 to $1,775 depending on the model.

Sol costs $5 per million input tokens and $30 per million output tokens. Terra is exactly half at $2.50/$15, and Luna is one-fifth at $1/$6. The ratios hold across workloads, so an agent task costing $0.63 on Sol costs $0.315 on Terra and $0.126 on Luna.

Yes. All five GPT-5.6 models are reasoning models, and thinking tokens are billed at the output rate even though they never appear in the response, so a task returning 8,000 visible tokens can bill 14,000 or more. Because output costs 6x input on every 5.6 model, reasoning effort settings often move your bill more than prompt length does.

The discounts multiply. Cached input bills at roughly 10% of base and the Batch API takes 50% off everything, so cached tokens inside a batch job cost about 5% of list price. A reasoning-heavy agent task on Sol drops from $1.17 at list price to $0.315 with 80% cache hits plus batching, a 73% reduction. The trade-off is that batch jobs complete within a 24-hour window rather than in real time.

On cost, yes. At $0.75/$4.50 per million tokens, GPT-5.4 mini prices a standard chat exchange at $0.002662, exactly 25% below GPT-5.6 Luna’s $0.00355. Luna adds reasoning ability and a roughly 1M-token context window, but mini remains OpenAI’s cheapest option for high-volume simple tasks.

GPT-5.5, released in April 2026, doubled flagship pricing to $5/$30 and shipped without a mid-tier, then the GPT-5.6 launch on July 9, 2026 restored the ladder with Terra at $2.50/$15 and Luna at $1/$6. Teams that hardcoded a model ID in April absorbed a 2x cost increase overnight. Keeping model choice in configuration and re-benchmarking on each release avoids repeating that.