Calculate OpenAI API costs for GPT-5.6 Sol, Terra, and Luna with live per-request pricing, cache and batch discounts, and worked examples per workload.
| Model | $ / 1M in · out | Cost / request | Cost / 1K requests | Cost / month |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | $5.00 · $30.00 | $0.017750 | $17.750 | $1775.00 |
| GPT-5.6 Terra OpenAI | $2.50 · $15.00 | $0.008875 | $8.875 | $887.50 |
| GPT-5.6 Luna OpenAI | $1.00 · $6.00 | $0.003550 | $3.550 | $355.00 |
| GPT-5.4 mini OpenAI | $0.75 · $4.50 | $0.002662 | $2.663 | $266.25 |
| GPT-5.4 nano OpenAI | $0.20 · $1.25 | $0.000735 | $0.735 | $73.50 |
| Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3× | $10.00 · $50.00 | $0.031125 | $31.125 | $3112.50 |
| Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3× | $5.00 · $25.00 | $0.015562 | $15.562 | $1556.25 |
| Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled | $2.00 · $10.00 | $0.009338 | $9.338 | $933.75 |
| Claude Haiku 4.5 Anthropic (Claude) | $1.00 · $5.00 | $0.003112 | $3.112 | $311.25 |
| Gemini 3.1 Pro (Preview) Google (Gemini API) | $2.00 · $12.00 | $0.007100 | $7.100 | $710.00 |
| Gemini 3.5 Flash Google (Gemini API) | $1.50 · $9.00 | $0.005325 | $5.325 | $532.50 |
| Gemini 2.5 Flash-Lite Google (Gemini API) | $0.10 · $0.40 | $0.000255 | $0.255 | $25.50 |
| DeepSeek V4 Pro DeepSeek | $0.43 · $0.87 | $0.000654 | $0.654 | $65.43 |
| DeepSeek V4 Flash DeepSeek | $0.14 · $0.28 | $0.000211 | $0.211 | $21.14 |
| Mistral Medium 3.5 Mistral | $1.50 · $7.50 | $0.005250 | $5.250 | $525.00 |
| Mistral Large 3 Mistral | $0.50 · $1.50 | $0.001025 | $1.025 | $102.50 |
| Mistral Small 4 Mistral | $0.15 · $0.60 | $0.000382 | $0.382 | $38.25 |
| Grok 4.5 xAI (Grok) | $2.00 · $6.00 | $0.004250 | $4.250 | $425.00 |
| Grok 4.3 xAI (Grok) | $1.25 · $2.50 | $0.001975 | $1.975 | $197.50 |
| Llama 4 Scout Meta Llama (hosted) | $0.10 · $0.30 | $0.000250 | $0.250 | $25.00 |
| Llama 4 Maverick Meta Llama (hosted) | $0.15 · $0.60 | $0.000450 | $0.450 | $45.00 |
| Amazon Nova Pro Amazon (Nova on Bedrock) | $0.80 · $3.20 | $0.002400 | $2.400 | $240.00 |
| Amazon Nova Lite Amazon (Nova on Bedrock) | $0.06 · $0.24 | $0.000180 | $0.180 | $18.00 |
| Amazon Nova Micro Amazon (Nova on Bedrock) | $0.04 · $0.14 | $0.000105 | $0.105 | $10.50 |
The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.
OpenAI’s GPT-5.6 family, launched July 9, 2026, prices its three general-purpose tiers in fixed ratios, with Sol at $5 input / $30 output per million tokens, Terra at exactly half ($2.50/$15), and Luna at one-fifth ($1/$6). Because the ratios are constant, per-request costs scale predictably across the three workload profiles the calculator models.
At 100,000 chat requests a month, that is $1,775 on Sol, $887.50 on Terra, or $355 on Luna. GPT-5.4 mini ($0.75/$4.50) still undercuts the whole 5.6 ladder at $0.002662 per chat exchange, exactly 25% below Luna, and cheaper rows exist outside OpenAI entirely, such as Grok 4.5 at $0.00425 and DeepSeek V4 Pro at $0.000654 for the same exchange.
Terra costs exactly 2.5x Luna on every line item, so the question is whether accuracy gains justify a 2.5x bill. At 1,000 agent tasks per month the gap is $189 ($315 versus $126), and at 10,000 RAG queries it is $101.70 ($169.50 versus $67.80). Run evals both ways. If Luna passes, the premium buys nothing, and if Terra prevents even a small share of failed agent runs that would be retried at full cost, it can pay for itself. Sol’s further 2x step over Terra follows the same logic at higher stakes.
All five GPT-5.6 models are reasoning models, and thinking tokens bill at the output rate without appearing in the response. The agent-task figures above assume a 1.75x reasoning multiplier, so 8,000 visible output tokens bill as 14,000. On Sol, that output line is $0.42 of the $0.63 total, two-thirds of the cost of a task whose input is nearly 19x larger than its visible output. Because output costs 6x input across the entire ladder, tuning reasoning effort down on tasks that don’t need deep deliberation usually saves more than trimming prompts.
OpenAI changed its price structure twice in three months this year. GPT-5.5, released in April 2026, doubled flagship pricing to $5/$30 and shipped with no mid-tier, leaving GPT-5.4 mini ($0.75/$4.50) and nano ($0.20/$1.25) as the only budget options. The GPT-5.6 launch on July 9 kept the $5/$30 flagship rate for Sol but restored the middle of the ladder with Terra and Luna. Teams that hardcoded a flagship model ID in April absorbed a 2x cost increase overnight, while teams that route the model name through configuration re-ran their evals and switched. The practical lesson is to treat the model ID as a config value, re-benchmark on every release, and let current prices drive the decision rather than a constant committed to your codebase six months ago.
The two discounts multiply. Cached input bills at roughly 10% of base, and the Batch API halves everything, so a cached input token inside a batch job costs about 5% of list price. Work it through the agent task on Sol. At list price with no caching, 150K input tokens cost $0.75 and 14K billed output tokens cost $0.42, or $1.17 per task. An 80% cache-hit rate cuts that to $0.63, and running it through the Batch API cuts it to $0.315, a 73% total reduction. Push cache hits above 90% and the combined saving exceeds 75%. The constraints are concrete. Batch jobs complete within a 24-hour window rather than in real time, and caching only hits when prompt prefixes are byte-stable across requests, so put system prompts and tool definitions first and variable content last.
For high-volume simple chat, GPT-5.4 mini at $0.002662 per exchange is the cheapest OpenAI option, with Luna the step up when reasoning or long context matters. For RAG, Luna’s $0.00678 per query covers most retrieval-grounded answering, and Terra’s $0.01695 is the upgrade when synthesis across many retrieved chunks needs to be reliable. For agents, a common pattern is Luna for routine steps with escalation to Sol for planning. At $0.126 versus $0.63 per task, an 80/20 split averages $0.227 per task, roughly a third of running everything on Sol. All five 5.6 models share context windows of about 1M tokens at a flat rate, so context length never forces a price-tier change the way it does on Gemini 3.1 Pro’s two-tier pricing.