Price usage-based plans from real token costs. Credit pack math, seat-plus-usage hybrids, overage design, and the metering AI billing actually requires.
| Model | $ / 1M in · out | Cost / request | Cost / 1K requests | Cost / month |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | $5.00 · $30.00 | $0.017750 | $17.750 | $1775.00 |
| GPT-5.6 Terra OpenAI | $2.50 · $15.00 | $0.008875 | $8.875 | $887.50 |
| GPT-5.6 Luna OpenAI | $1.00 · $6.00 | $0.003550 | $3.550 | $355.00 |
| GPT-5.4 mini OpenAI | $0.75 · $4.50 | $0.002662 | $2.663 | $266.25 |
| GPT-5.4 nano OpenAI | $0.20 · $1.25 | $0.000735 | $0.735 | $73.50 |
| Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3× | $10.00 · $50.00 | $0.031125 | $31.125 | $3112.50 |
| Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3× | $5.00 · $25.00 | $0.015562 | $15.562 | $1556.25 |
| Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled | $2.00 · $10.00 | $0.009338 | $9.338 | $933.75 |
| Claude Haiku 4.5 Anthropic (Claude) | $1.00 · $5.00 | $0.003112 | $3.112 | $311.25 |
| Gemini 3.1 Pro (Preview) Google (Gemini API) | $2.00 · $12.00 | $0.007100 | $7.100 | $710.00 |
| Gemini 3.5 Flash Google (Gemini API) | $1.50 · $9.00 | $0.005325 | $5.325 | $532.50 |
| Gemini 2.5 Flash-Lite Google (Gemini API) | $0.10 · $0.40 | $0.000255 | $0.255 | $25.50 |
| DeepSeek V4 Pro DeepSeek | $0.43 · $0.87 | $0.000654 | $0.654 | $65.43 |
| DeepSeek V4 Flash DeepSeek | $0.14 · $0.28 | $0.000211 | $0.211 | $21.14 |
| Mistral Medium 3.5 Mistral | $1.50 · $7.50 | $0.005250 | $5.250 | $525.00 |
| Mistral Large 3 Mistral | $0.50 · $1.50 | $0.001025 | $1.025 | $102.50 |
| Mistral Small 4 Mistral | $0.15 · $0.60 | $0.000382 | $0.382 | $38.25 |
| Grok 4.5 xAI (Grok) | $2.00 · $6.00 | $0.004250 | $4.250 | $425.00 |
| Grok 4.3 xAI (Grok) | $1.25 · $2.50 | $0.001975 | $1.975 | $197.50 |
| Llama 4 Scout Meta Llama (hosted) | $0.10 · $0.30 | $0.000250 | $0.250 | $25.00 |
| Llama 4 Maverick Meta Llama (hosted) | $0.15 · $0.60 | $0.000450 | $0.450 | $45.00 |
| Amazon Nova Pro Amazon (Nova on Bedrock) | $0.80 · $3.20 | $0.002400 | $2.400 | $240.00 |
| Amazon Nova Lite Amazon (Nova on Bedrock) | $0.06 · $0.24 | $0.000180 | $0.180 | $18.00 |
| Amazon Nova Micro Amazon (Nova on Bedrock) | $0.04 · $0.14 | $0.000105 | $0.105 | $10.50 |
The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.
Gross margin counts API cost only. Hosting, support, and payment fees come out of the same price. Healthy AI products keep API costs to 15–30% of revenue, and a common rule of thumb is charging 3.5–5× your token cost. The usage-growth rows show what happens to margin when your users get more active without paying more.
Charging for usage means metering it. Kelviq bundles usage metering, entitlements, and merchant-of-record billing for AI products into one transaction fee.
Usage-based pricing is not one model but three, and most products end up combining them, and 92% of AI software companies mix subscription and usage pricing.
Customers pay exactly for what they consume, per request or per token. This is the strongest margin protection, since cost and revenue move together, but the weakest bill predictability, which enterprise buyers in particular resist. It fits API-style products where the customer is technical and expects a meter.
The most common pricing model among AI wrapper products. Customers prepay for a bundle of credits, and each action consumes a defined number. Packs bound your exposure, simplify the buying decision, and introduce breakage, meaning revenue from credits never consumed. Cursor Pro illustrates the extreme. Its $20/month for roughly $20 of usage is about 0% margin on consumed usage, with margin coming from unused allowances.
A base subscription covers an included allowance. Metered usage or credit top-ups cover the rest. This is the dominant pattern across software because it gives finance teams a predictable floor and protects the vendor’s tail risk.
Credits only work if they are anchored to a real cost unit. Here is a worked example using the live anchors above.
Define 1 credit = 1 chat request (1,000 input / 500 output tokens, 50% cached) on Claude Sonnet 5, which costs $0.006225. A 1,000-credit pack therefore costs you $6.23. Applying the standard rule of thumb of charging 3.5-5× token cost for a 70-80% gross margin, the pack should sell for $22 to $31. A $25 pack is exactly 4.0× cost, a 75.1% gross margin. The common $20 and $30 price points land at 68.9% and 79.2%.
For multi-model products, weight credits by relative cost rather than averaging. A Claude Fable 5 request at $0.031125 is exactly 5× a Sonnet 5 request, so charging 5 credits for it keeps every route at the same margin. A timing note applies here. Sonnet 5’s figure is intro pricing that rises about 50% on September 1, 2026, so credit definitions pegged to it need a re-anchor then.
An allowance needs a policy for what happens at its edge. Hard caps (usage stops) suit free tiers, where every marginal request is pure cost. Soft caps with overage billing suit paid tiers, but the overage rate must hold your multiple. If the base pack works out to $0.025 per credit, overage priced below that means your heaviest users buy their cheapest usage exactly where your cost risk concentrates.
The margin panel above shows margin at 2× and 5× usage for this reason. On a $20 plan backed by Sonnet 5, margin falls from 90.7% at 300 requests to 53.3% at 1,500. Wherever that curve crosses your minimum acceptable margin is where the included allowance should end.
None of the above is billable without per-request metering of model, input tokens, output tokens, cached share, account, and timestamp. Aggregate monthly totals cannot settle a disputed invoice or tell you which plan tier is losing money. This is the infrastructure layer tools like Kelviq handle, mapping per-request usage onto credit and hybrid plans, though the same need can be met with in-house event pipelines if billing complexity stays low.
Usage components earn their complexity when per-user cost is genuinely variable. Agent workloads are the clearest case. Tasks span $0.0257 (DeepSeek V4 Pro) to $1.27 (Claude Fable 5), and a heavy user running several flagship sessions daily can burn $200-500 per month. No flat seat survives that unmetered. Flat seats remain the right call when usage is uniform and API costs sit inside the healthy 15-30% of revenue band. The honest default is to start flat, meter from day one, and add usage pricing when the data shows the spread.
Base per-token prices refresh weekly from the OpenRouter model feed (last refresh 2026-07-11). Everything the feed can’t see, such as caching mechanics, batch discounts, long-context tiers, and announced price changes, is verified by hand against each provider’s official API pricing page.