Claude API Cost Calculator for Fable, Opus & Sonnet

Claude API pricing calculator for Fable 5, Opus 4.8, Sonnet 5, and Haiku 4.5, with the September 1 Sonnet price rise and cache-write costs modeled.

Model$ / 1M in · outCost / requestCost / 1K requestsCost / month
GPT-5.6 Sol OpenAI$5.00 · $30.00$0.017750$17.750$1775.00
GPT-5.6 Terra OpenAI$2.50 · $15.00$0.008875$8.875$887.50
GPT-5.6 Luna OpenAI$1.00 · $6.00$0.003550$3.550$355.00
GPT-5.4 mini OpenAI$0.75 · $4.50$0.002662$2.663$266.25
GPT-5.4 nano OpenAI$0.20 · $1.25$0.000735$0.735$73.50
Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3×$10.00 · $50.00$0.031125$31.125$3112.50
Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3×$5.00 · $25.00$0.015562$15.562$1556.25
Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled$2.00 · $10.00$0.009338$9.338$933.75
Claude Haiku 4.5 Anthropic (Claude)$1.00 · $5.00$0.003112$3.112$311.25
Gemini 3.1 Pro (Preview) Google (Gemini API)$2.00 · $12.00$0.007100$7.100$710.00
Gemini 3.5 Flash Google (Gemini API)$1.50 · $9.00$0.005325$5.325$532.50
Gemini 2.5 Flash-Lite Google (Gemini API)$0.10 · $0.40$0.000255$0.255$25.50
DeepSeek V4 Pro DeepSeek$0.43 · $0.87$0.000654$0.654$65.43
DeepSeek V4 Flash DeepSeek$0.14 · $0.28$0.000211$0.211$21.14
Mistral Medium 3.5 Mistral$1.50 · $7.50$0.005250$5.250$525.00
Mistral Large 3 Mistral$0.50 · $1.50$0.001025$1.025$102.50
Mistral Small 4 Mistral$0.15 · $0.60$0.000382$0.382$38.25
Grok 4.5 xAI (Grok)$2.00 · $6.00$0.004250$4.250$425.00
Grok 4.3 xAI (Grok)$1.25 · $2.50$0.001975$1.975$197.50
Llama 4 Scout Meta Llama (hosted)$0.10 · $0.30$0.000250$0.250$25.00
Llama 4 Maverick Meta Llama (hosted)$0.15 · $0.60$0.000450$0.450$45.00
Amazon Nova Pro Amazon (Nova on Bedrock)$0.80 · $3.20$0.002400$2.400$240.00
Amazon Nova Lite Amazon (Nova on Bedrock)$0.06 · $0.24$0.000180$0.180$18.00
Amazon Nova Micro Amazon (Nova on Bedrock)$0.04 · $0.14$0.000105$0.105$10.50

The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.

How Anthropic (Claude) pricing works

  • Prompt caching: Cache reads cost 10% of base input, while cache WRITES cost 1.25x (5-min TTL) or 2x (1-hour TTL)
  • Batch API: −50% on asynchronous workloads.
  • Opus 4.7+, Sonnet 5, and Fable 5 use a new tokenizer producing up to roughly 30% more tokens for the same text, so cross-provider per-token comparisons understate Claude cost slightly.
  • 1M-token context at a flat price with no long-context surcharge.
  • Source: official pricing page, details last verified 2026-07-11.

Sonnet 5’s intro price ends September 1, so model your budget at $3/$15

Anthropic’s published pricing table lists Sonnet 5 at $2 input / $10 output per million tokens as an introductory rate, rising to $3/$15 on September 1, 2026. Every Sonnet figure in the calculator above reflects the intro rate, so any forecast that extends past August needs a 50% uplift.

WorkloadIntro ($2/$10)From Sep 1 ($3/$15)
Chat exchange (1,000 in / 500 out, 50% cached)$0.006225~$0.00934
RAG query (6,000 in / 400 out, 30% cached)$0.01321~$0.01982
Agent task (150K in / 8K out, 80% cached)$0.254$0.381

At 100,000 chat requests a month, that is $622.50 today and $933.75 from September. The decision rule is simple. If Sonnet 5 clears your unit economics at $3/$15, the intro period is a bonus. If your margins only work at $2/$10, the model choice needs revisiting before the rise, not after it lands on an invoice.

Cache writes cost money, so caching pays only on reuse

Claude’s prompt caching bills in both directions. Reads cost 10% of the input rate, but writes carry a premium of 1.25x base for the 5-minute TTL or 2x for the 1-hour TTL. A one-shot prompt that writes a cache and never reads it costs 25-100% more than skipping caching entirely. The break-even arrives quickly, though. Each read within the TTL saves 0.9x the base rate, so a 5-minute cache pays for itself on the first reuse and a 1-hour cache on the second.

Agent workloads are the clear win case. The Fable 5 agent-task figure of $1.27 assumes 80% of the 150K-token context hits cache. The identical task with no caching bills $2.20 ($1.50 of input plus $0.70 for 14,000 billed output tokens), so caching cuts 42% even after write premiums. Note that the calculator’s workload figures include amortized cache-write costs, which is why they run slightly above a naive estimate that only applies the 90% read discount.

The tokenizer caveat makes per-token rates understate Claude’s cost

Opus 4.7 and later, Sonnet 5, and Fable 5 use a new tokenizer that produces up to roughly 30% more tokens than previous Claude models for the same text. That makes per-token rate comparisons across providers misleading. Sonnet 5’s $2 input rate reads identically to Gemini 3.1 Pro’s $2, but the same document can tokenize substantially larger on Claude, and output inherits the same inflation. The workload figures in this calculator compare models at fixed token counts, so when estimating against your own corpus, tokenize a representative sample with each provider’s counter before trusting any per-million-token rate.

Fable 5, Opus 4.8, or Sonnet 5

The ladder steps are steep but regular. Fable 5 ($10/$50) costs exactly 2x Opus 4.8 ($5/$25), which runs 2.5x Sonnet 5’s introductory rate (1.67x after September 1). Per chat exchange that is $0.031125, $0.015562, and $0.006225, and per agent task it works out to $1.27, $0.635, and $0.254. Opus 4.8 is the sensible default flagship, at half of Fable 5 on every workload, with Fable 5 reserved for tasks where evals show the premium model materially changing outcomes. One rung doesn’t appear in the calculator. Opus 4.8 has a fast mode priced at $10/$50, the same nominal rate as Fable 5, which purchases latency rather than capability. All of these models offer 1M-token context at a flat price, with no surcharge tier above 200K tokens of the kind Gemini 3.1 Pro applies.

Haiku 4.5 for volume

Haiku 4.5 at $1/$5 prices a chat exchange at $0.003112, so one million exchanges a month costs $3,112, versus $6,225 on Sonnet 5 intro and $31,125 on Fable 5. Its agent-task cost of $0.097 is 13x below Fable 5’s $1.27, and the 50% batch discount stacks on top, taking asynchronous Haiku agent tasks to $0.0485 each. For classification, routing, extraction, and other high-volume steps inside a larger pipeline, running Haiku by default and escalating to Opus 4.8 on failure keeps blended cost close to the Haiku floor while preserving a quality backstop.

Frequently asked questions

A standard chat exchange (1,000 input tokens, 500 output tokens, 50% cache hits) costs $0.031125 on Claude Fable 5, $0.015562 on Opus 4.8, $0.006225 on Sonnet 5 at its introductory rate, and $0.003112 on Haiku 4.5. Those figures include prompt cache-write costs, which Anthropic bills on top of discounted cache reads.

Anthropic’s official pricing table lists Sonnet 5’s $2/$10 per-million-token rate as introductory, rising to $3/$15 on September 1, 2026. That is a 50% increase, taking a typical chat exchange from $0.006225 to about $0.00934 and an agent task from $0.254 to $0.381. Budget forecasts past August should use the $3/$15 rate.

Cache reads cost 10% of the normal input rate, but writing to the cache costs extra, at 1.25x base for the 5-minute TTL or 2x for the 1-hour TTL. Each read within the TTL saves 0.9x the base rate, so a 5-minute cache pays for itself after one reuse and a 1-hour cache after two. One-shot prompts that never reuse the cache cost more than not caching at all.

It can. Opus 4.7 and later, Sonnet 5, and Fable 5 use a new tokenizer that produces up to about 30% more tokens than previous Claude models for the same text. Comparing per-token rates across providers therefore understates Claude’s cost per document. Comparing per-request costs at measured token counts is more reliable.

Fable 5 costs exactly double Opus 4.8, at $10/$50 versus $5/$25 per million tokens, or $1.27 versus $0.635 for a large cached agent task. Anthropic also offers an Opus 4.8 fast mode at $10/$50, which matches Fable 5’s rates while buying latency rather than capability.

Yes, batch processing takes 50% off both input and output tokens and stacks with prompt caching. A Sonnet 5 agent task at the introductory rate drops from $0.254 to $0.127 when batched, and a Haiku 4.5 agent task from $0.097 to $0.0485. Batch jobs run asynchronously rather than in real time.