Usage-Based Pricing Calculator for AI & SaaS

Price usage-based plans from real token costs. Credit pack math, seat-plus-usage hybrids, overage design, and the metering AI billing actually requires.

Model$ / 1M in · outCost / requestCost / 1K requestsCost / month
GPT-5.6 Sol OpenAI$5.00 · $30.00$0.017750$17.750$1775.00
GPT-5.6 Terra OpenAI$2.50 · $15.00$0.008875$8.875$887.50
GPT-5.6 Luna OpenAI$1.00 · $6.00$0.003550$3.550$355.00
GPT-5.4 mini OpenAI$0.75 · $4.50$0.002662$2.663$266.25
GPT-5.4 nano OpenAI$0.20 · $1.25$0.000735$0.735$73.50
Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3×$10.00 · $50.00$0.031125$31.125$3112.50
Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3×$5.00 · $25.00$0.015562$15.562$1556.25
Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled$2.00 · $10.00$0.009338$9.338$933.75
Claude Haiku 4.5 Anthropic (Claude)$1.00 · $5.00$0.003112$3.112$311.25
Gemini 3.1 Pro (Preview) Google (Gemini API)$2.00 · $12.00$0.007100$7.100$710.00
Gemini 3.5 Flash Google (Gemini API)$1.50 · $9.00$0.005325$5.325$532.50
Gemini 2.5 Flash-Lite Google (Gemini API)$0.10 · $0.40$0.000255$0.255$25.50
DeepSeek V4 Pro DeepSeek$0.43 · $0.87$0.000654$0.654$65.43
DeepSeek V4 Flash DeepSeek$0.14 · $0.28$0.000211$0.211$21.14
Mistral Medium 3.5 Mistral$1.50 · $7.50$0.005250$5.250$525.00
Mistral Large 3 Mistral$0.50 · $1.50$0.001025$1.025$102.50
Mistral Small 4 Mistral$0.15 · $0.60$0.000382$0.382$38.25
Grok 4.5 xAI (Grok)$2.00 · $6.00$0.004250$4.250$425.00
Grok 4.3 xAI (Grok)$1.25 · $2.50$0.001975$1.975$197.50
Llama 4 Scout Meta Llama (hosted)$0.10 · $0.30$0.000250$0.250$25.00
Llama 4 Maverick Meta Llama (hosted)$0.15 · $0.60$0.000450$0.450$45.00
Amazon Nova Pro Amazon (Nova on Bedrock)$0.80 · $3.20$0.002400$2.400$240.00
Amazon Nova Lite Amazon (Nova on Bedrock)$0.06 · $0.24$0.000180$0.180$18.00
Amazon Nova Micro Amazon (Nova on Bedrock)$0.04 · $0.14$0.000105$0.105$10.50

The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.

What should you charge?

API cost / user / month $2.80
Gross margin at your price 86.0%
Break-even price $2.80
Price for 70% margin $9.34
Price for 80% margin $14.01
Margin if usage doubles 72.0%
Margin at 5× usage 30.0%

Gross margin counts API cost only. Hosting, support, and payment fees come out of the same price. Healthy AI products keep API costs to 15–30% of revenue, and a common rule of thumb is charging 3.5–5× your token cost. The usage-growth rows show what happens to margin when your users get more active without paying more.

Charging for usage means metering it. Kelviq bundles usage metering, entitlements, and merchant-of-record billing for AI products into one transaction fee.

The three shapes of usage pricing

Usage-based pricing is not one model but three, and most products end up combining them, and 92% of AI software companies mix subscription and usage pricing.

Pure metered usage

Customers pay exactly for what they consume, per request or per token. This is the strongest margin protection, since cost and revenue move together, but the weakest bill predictability, which enterprise buyers in particular resist. It fits API-style products where the customer is technical and expects a meter.

Credit packs

The most common pricing model among AI wrapper products. Customers prepay for a bundle of credits, and each action consumes a defined number. Packs bound your exposure, simplify the buying decision, and introduce breakage, meaning revenue from credits never consumed. Cursor Pro illustrates the extreme. Its $20/month for roughly $20 of usage is about 0% margin on consumed usage, with margin coming from unused allowances.

Hybrid seat plus usage

A base subscription covers an included allowance. Metered usage or credit top-ups cover the rest. This is the dominant pattern across software because it gives finance teams a predictable floor and protects the vendor’s tail risk.

Setting a credit price from token cost

Credits only work if they are anchored to a real cost unit. Here is a worked example using the live anchors above.

Define 1 credit = 1 chat request (1,000 input / 500 output tokens, 50% cached) on Claude Sonnet 5, which costs $0.006225. A 1,000-credit pack therefore costs you $6.23. Applying the standard rule of thumb of charging 3.5-5× token cost for a 70-80% gross margin, the pack should sell for $22 to $31. A $25 pack is exactly 4.0× cost, a 75.1% gross margin. The common $20 and $30 price points land at 68.9% and 79.2%.

For multi-model products, weight credits by relative cost rather than averaging. A Claude Fable 5 request at $0.031125 is exactly 5× a Sonnet 5 request, so charging 5 credits for it keeps every route at the same margin. A timing note applies here. Sonnet 5’s figure is intro pricing that rises about 50% on September 1, 2026, so credit definitions pegged to it need a re-anchor then.

Caps and overage design

An allowance needs a policy for what happens at its edge. Hard caps (usage stops) suit free tiers, where every marginal request is pure cost. Soft caps with overage billing suit paid tiers, but the overage rate must hold your multiple. If the base pack works out to $0.025 per credit, overage priced below that means your heaviest users buy their cheapest usage exactly where your cost risk concentrates.

The margin panel above shows margin at 2× and 5× usage for this reason. On a $20 plan backed by Sonnet 5, margin falls from 90.7% at 300 requests to 53.3% at 1,500. Wherever that curve crosses your minimum acceptable margin is where the included allowance should end.

Metering, the unglamorous requirement

None of the above is billable without per-request metering of model, input tokens, output tokens, cached share, account, and timestamp. Aggregate monthly totals cannot settle a disputed invoice or tell you which plan tier is losing money. This is the infrastructure layer tools like Kelviq handle, mapping per-request usage onto credit and hybrid plans, though the same need can be met with in-house event pipelines if billing complexity stays low.

When usage pricing beats seats

Usage components earn their complexity when per-user cost is genuinely variable. Agent workloads are the clearest case. Tasks span $0.0257 (DeepSeek V4 Pro) to $1.27 (Claude Fable 5), and a heavy user running several flagship sessions daily can burn $200-500 per month. No flat seat survives that unmetered. Flat seats remain the right call when usage is uniform and API costs sit inside the healthy 15-30% of revenue band. The honest default is to start flat, meter from day one, and add usage pricing when the data shows the spread.

Price sources & freshness

Base per-token prices refresh weekly from the OpenRouter model feed (last refresh 2026-07-11). Everything the feed can’t see, such as caching mechanics, batch discounts, long-context tiers, and announced price changes, is verified by hand against each provider’s official API pricing page.

  • OpenAI official pricing, details last verified 2026-07-11 (The GPT-5.6 family (Sol/Terra/Luna) launched Jul 9, 2026. Its three-tier naming replaces the mini/nano ladder at the top end.)
  • Anthropic (Claude) official pricing, details last verified 2026-07-11 (Opus 4.7+, Sonnet 5, and Fable 5 use a new tokenizer producing up to roughly 30% more tokens for the same text, so cross-provider per-token comparisons understate Claude cost slightly.)
  • Google (Gemini API) official pricing, details last verified 2026-07-11 (Gemini 3.1 Pro doubles input pricing above 200K input tokens, the only two-tier context pricing in this comparison.)
  • DeepSeek official pricing, details last verified 2026-07-11 (Prices below are DeepSeek's first-party API. OpenRouter routes some DeepSeek models through discounted hosts at different prices.)
  • Mistral official pricing, details last verified 2026-07-11 (Verify against the API pricing page, since Mistral's marketing pages have shown stale prices for retired models.)
  • xAI (Grok) official pricing, details last verified 2026-07-11
  • Meta Llama (hosted) official pricing, details last verified 2026-07-11 (These are open-weight models, so the same model costs different amounts on different hosts such as Groq, DeepInfra, and Bedrock. Prices shown are OpenRouter's routed rate, indicative rather than contractual.)
  • Amazon (Nova on Bedrock) official pricing, details last verified 2026-07-11 (Bedrock also hosts Anthropic, Mistral, and DeepSeek models, with Claude at first-party prices and DeepSeek at different ones. Only Amazon's own Nova line is listed here.)

Frequently asked questions

Usage-based pricing charges customers in proportion to what they consume (per request, per token, per credit, or per task) instead of a flat seat fee. It rarely appears alone. Some 92% of AI software companies mix subscription and usage pricing, typically a base seat with an included allowance plus metered usage beyond it.

The three common shapes are pure metered usage (pay exactly for what you use), prepaid credit packs, and hybrid seat-plus-usage plans. Credit packs are the most common model among AI products, while hybrids dominate the broader software market. The right shape depends on how variable your per-user cost is and how much bill predictability customers demand.

Anchor the credit to a real unit of cost, then apply a 3.5-5x multiple. If one credit equals one chat request on Claude Sonnet 5 at $0.006225, a 1,000-credit pack costs you $6.23. Pricing it between $22 and $31 sits inside the 3.5-5x band, and a $25 pack is exactly 4x cost at a 75.1% gross margin. The common $20 and $30 price points yield 68.9% and 79.2% margins respectively.

Breakage is revenue from credits or allowances that customers pay for but never consume. Cursor Pro, for example, sells $20 per month of roughly $20 in usage, near 0% margin on consumed usage, with the real margin coming from unused allowances. Breakage improves realized margin but shrinks as users learn to exhaust their credits.

A plan includes an allowance, and usage beyond it is either blocked (hard cap) or billed at a per-unit overage rate (soft cap). Overage rates should preserve your target multiple, staying at or above the effective per-credit rate of the base pack. Otherwise heavy users buy their cheapest usage precisely where your risk is highest. Hard caps suit free tiers, while paid tiers usually pair soft caps with usage alerts.

Usage pricing wins when cost per user varies widely (agent products span $0.0257 to $1.27 per task depending on model) or when a heavy tail of users would be unprofitable at any flat price. Flat seats work when usage is uniform and API costs stay in the healthy 15-30% of revenue range. Many teams start flat and add usage components once real consumption data shows the spread.

At minimum you need per-request records of model used, input and output tokens, cached-token share, the account or user responsible, and a timestamp. Billing accuracy depends on this granularity, because you cannot reconstruct a disputed invoice from aggregate monthly totals. The same records also feed margin analysis and abuse detection.