Claude API pricing calculator for Fable 5, Opus 4.8, Sonnet 5, and Haiku 4.5, with the September 1 Sonnet price rise and cache-write costs modeled.
| Model | $ / 1M in · out | Cost / request | Cost / 1K requests | Cost / month |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | $5.00 · $30.00 | $0.017750 | $17.750 | $1775.00 |
| GPT-5.6 Terra OpenAI | $2.50 · $15.00 | $0.008875 | $8.875 | $887.50 |
| GPT-5.6 Luna OpenAI | $1.00 · $6.00 | $0.003550 | $3.550 | $355.00 |
| GPT-5.4 mini OpenAI | $0.75 · $4.50 | $0.002662 | $2.663 | $266.25 |
| GPT-5.4 nano OpenAI | $0.20 · $1.25 | $0.000735 | $0.735 | $73.50 |
| Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3× | $10.00 · $50.00 | $0.031125 | $31.125 | $3112.50 |
| Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3× | $5.00 · $25.00 | $0.015562 | $15.562 | $1556.25 |
| Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled | $2.00 · $10.00 | $0.009338 | $9.338 | $933.75 |
| Claude Haiku 4.5 Anthropic (Claude) | $1.00 · $5.00 | $0.003112 | $3.112 | $311.25 |
| Gemini 3.1 Pro (Preview) Google (Gemini API) | $2.00 · $12.00 | $0.007100 | $7.100 | $710.00 |
| Gemini 3.5 Flash Google (Gemini API) | $1.50 · $9.00 | $0.005325 | $5.325 | $532.50 |
| Gemini 2.5 Flash-Lite Google (Gemini API) | $0.10 · $0.40 | $0.000255 | $0.255 | $25.50 |
| DeepSeek V4 Pro DeepSeek | $0.43 · $0.87 | $0.000654 | $0.654 | $65.43 |
| DeepSeek V4 Flash DeepSeek | $0.14 · $0.28 | $0.000211 | $0.211 | $21.14 |
| Mistral Medium 3.5 Mistral | $1.50 · $7.50 | $0.005250 | $5.250 | $525.00 |
| Mistral Large 3 Mistral | $0.50 · $1.50 | $0.001025 | $1.025 | $102.50 |
| Mistral Small 4 Mistral | $0.15 · $0.60 | $0.000382 | $0.382 | $38.25 |
| Grok 4.5 xAI (Grok) | $2.00 · $6.00 | $0.004250 | $4.250 | $425.00 |
| Grok 4.3 xAI (Grok) | $1.25 · $2.50 | $0.001975 | $1.975 | $197.50 |
| Llama 4 Scout Meta Llama (hosted) | $0.10 · $0.30 | $0.000250 | $0.250 | $25.00 |
| Llama 4 Maverick Meta Llama (hosted) | $0.15 · $0.60 | $0.000450 | $0.450 | $45.00 |
| Amazon Nova Pro Amazon (Nova on Bedrock) | $0.80 · $3.20 | $0.002400 | $2.400 | $240.00 |
| Amazon Nova Lite Amazon (Nova on Bedrock) | $0.06 · $0.24 | $0.000180 | $0.180 | $18.00 |
| Amazon Nova Micro Amazon (Nova on Bedrock) | $0.04 · $0.14 | $0.000105 | $0.105 | $10.50 |
The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.
Anthropic’s published pricing table lists Sonnet 5 at $2 input / $10 output per million tokens as an introductory rate, rising to $3/$15 on September 1, 2026. Every Sonnet figure in the calculator above reflects the intro rate, so any forecast that extends past August needs a 50% uplift.
| Workload | Intro ($2/$10) | From Sep 1 ($3/$15) |
|---|---|---|
| Chat exchange (1,000 in / 500 out, 50% cached) | $0.006225 | ~$0.00934 |
| RAG query (6,000 in / 400 out, 30% cached) | $0.01321 | ~$0.01982 |
| Agent task (150K in / 8K out, 80% cached) | $0.254 | $0.381 |
At 100,000 chat requests a month, that is $622.50 today and $933.75 from September. The decision rule is simple. If Sonnet 5 clears your unit economics at $3/$15, the intro period is a bonus. If your margins only work at $2/$10, the model choice needs revisiting before the rise, not after it lands on an invoice.
Claude’s prompt caching bills in both directions. Reads cost 10% of the input rate, but writes carry a premium of 1.25x base for the 5-minute TTL or 2x for the 1-hour TTL. A one-shot prompt that writes a cache and never reads it costs 25-100% more than skipping caching entirely. The break-even arrives quickly, though. Each read within the TTL saves 0.9x the base rate, so a 5-minute cache pays for itself on the first reuse and a 1-hour cache on the second.
Agent workloads are the clear win case. The Fable 5 agent-task figure of $1.27 assumes 80% of the 150K-token context hits cache. The identical task with no caching bills $2.20 ($1.50 of input plus $0.70 for 14,000 billed output tokens), so caching cuts 42% even after write premiums. Note that the calculator’s workload figures include amortized cache-write costs, which is why they run slightly above a naive estimate that only applies the 90% read discount.
Opus 4.7 and later, Sonnet 5, and Fable 5 use a new tokenizer that produces up to roughly 30% more tokens than previous Claude models for the same text. That makes per-token rate comparisons across providers misleading. Sonnet 5’s $2 input rate reads identically to Gemini 3.1 Pro’s $2, but the same document can tokenize substantially larger on Claude, and output inherits the same inflation. The workload figures in this calculator compare models at fixed token counts, so when estimating against your own corpus, tokenize a representative sample with each provider’s counter before trusting any per-million-token rate.
The ladder steps are steep but regular. Fable 5 ($10/$50) costs exactly 2x Opus 4.8 ($5/$25), which runs 2.5x Sonnet 5’s introductory rate (1.67x after September 1). Per chat exchange that is $0.031125, $0.015562, and $0.006225, and per agent task it works out to $1.27, $0.635, and $0.254. Opus 4.8 is the sensible default flagship, at half of Fable 5 on every workload, with Fable 5 reserved for tasks where evals show the premium model materially changing outcomes. One rung doesn’t appear in the calculator. Opus 4.8 has a fast mode priced at $10/$50, the same nominal rate as Fable 5, which purchases latency rather than capability. All of these models offer 1M-token context at a flat price, with no surcharge tier above 200K tokens of the kind Gemini 3.1 Pro applies.
Haiku 4.5 at $1/$5 prices a chat exchange at $0.003112, so one million exchanges a month costs $3,112, versus $6,225 on Sonnet 5 intro and $31,125 on Fable 5. Its agent-task cost of $0.097 is 13x below Fable 5’s $1.27, and the 50% batch discount stacks on top, taking asynchronous Haiku agent tasks to $0.0485 each. For classification, routing, extraction, and other high-volume steps inside a larger pipeline, running Haiku by default and escalating to Opus 4.8 on failure keeps blended cost close to the Haiku floor while preserving a quality backstop.