Calculate what each user costs you in LLM API spend. Light, medium, and heavy usage tiers, agent workloads, and why P95 usage matters more than the average.
| Model | $ / 1M in · out | Cost / request | Cost / 1K requests | Cost / month |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | $5.00 · $30.00 | $0.017750 | $17.750 | $1775.00 |
| GPT-5.6 Terra OpenAI | $2.50 · $15.00 | $0.008875 | $8.875 | $887.50 |
| GPT-5.6 Luna OpenAI | $1.00 · $6.00 | $0.003550 | $3.550 | $355.00 |
| GPT-5.4 mini OpenAI | $0.75 · $4.50 | $0.002662 | $2.663 | $266.25 |
| GPT-5.4 nano OpenAI | $0.20 · $1.25 | $0.000735 | $0.735 | $73.50 |
| Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3× | $10.00 · $50.00 | $0.031125 | $31.125 | $3112.50 |
| Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3× | $5.00 · $25.00 | $0.015562 | $15.562 | $1556.25 |
| Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled | $2.00 · $10.00 | $0.009338 | $9.338 | $933.75 |
| Claude Haiku 4.5 Anthropic (Claude) | $1.00 · $5.00 | $0.003112 | $3.112 | $311.25 |
| Gemini 3.1 Pro (Preview) Google (Gemini API) | $2.00 · $12.00 | $0.007100 | $7.100 | $710.00 |
| Gemini 3.5 Flash Google (Gemini API) | $1.50 · $9.00 | $0.005325 | $5.325 | $532.50 |
| Gemini 2.5 Flash-Lite Google (Gemini API) | $0.10 · $0.40 | $0.000255 | $0.255 | $25.50 |
| DeepSeek V4 Pro DeepSeek | $0.43 · $0.87 | $0.000654 | $0.654 | $65.43 |
| DeepSeek V4 Flash DeepSeek | $0.14 · $0.28 | $0.000211 | $0.211 | $21.14 |
| Mistral Medium 3.5 Mistral | $1.50 · $7.50 | $0.005250 | $5.250 | $525.00 |
| Mistral Large 3 Mistral | $0.50 · $1.50 | $0.001025 | $1.025 | $102.50 |
| Mistral Small 4 Mistral | $0.15 · $0.60 | $0.000382 | $0.382 | $38.25 |
| Grok 4.5 xAI (Grok) | $2.00 · $6.00 | $0.004250 | $4.250 | $425.00 |
| Grok 4.3 xAI (Grok) | $1.25 · $2.50 | $0.001975 | $1.975 | $197.50 |
| Llama 4 Scout Meta Llama (hosted) | $0.10 · $0.30 | $0.000250 | $0.250 | $25.00 |
| Llama 4 Maverick Meta Llama (hosted) | $0.15 · $0.60 | $0.000450 | $0.450 | $45.00 |
| Amazon Nova Pro Amazon (Nova on Bedrock) | $0.80 · $3.20 | $0.002400 | $2.400 | $240.00 |
| Amazon Nova Lite Amazon (Nova on Bedrock) | $0.06 · $0.24 | $0.000180 | $0.180 | $18.00 |
| Amazon Nova Micro Amazon (Nova on Bedrock) | $0.04 · $0.14 | $0.000105 | $0.105 | $10.50 |
The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.
Gross margin counts API cost only. Hosting, support, and payment fees come out of the same price. Healthy AI products keep API costs to 15–30% of revenue, and a common rule of thumb is charging 3.5–5× your token cost. The usage-growth rows show what happens to margin when your users get more active without paying more.
Charging for usage means metering it. Kelviq bundles usage metering, entitlements, and merchant-of-record billing for AI products into one transaction fee.
The formula is short. Cost per user = cost per request × requests per user per month. The first factor is knowable to four decimal places, and the calculator above computes it live for 24 models. The second factor is the one founders get wrong, because before launch nobody knows how often users will actually hit the product, and after launch the average hides a long tail.
Everything below uses the standard chat anchor (1,000 input / 500 output tokens per request, 50% cached), priced July 11, 2026.
A practical planning split is 100 / 300 / 1,000 requests per user per month.
| Model | Light (100/mo) | Medium (300/mo) | Heavy (1,000/mo) |
|---|---|---|---|
| Claude Fable 5 | $3.11 | $9.34 | $31.13 |
| GPT-5.6 Sol | $1.78 | $5.33 | $17.75 |
| Claude Sonnet 5 | $0.62 | $1.87 | $6.23 |
| GPT-5.6 Luna | $0.36 | $1.07 | $3.55 |
| Claude Haiku 4.5 | $0.31 | $0.93 | $3.11 |
| GPT-5.4 mini | $0.27 | $0.80 | $2.66 |
| DeepSeek V4 Pro | $0.07 | $0.20 | $0.65 |
The benchmark that turns these into decisions is that healthy AI products keep API costs to 15-30% of revenue. On a $20 seat, that is a $3-6 monthly cost budget per user. Sonnet 5 sits comfortably inside it at medium usage ($1.87), breaches it with heavy users ($6.23), and Fable 5 blows through it for anyone above light usage.
For agent workloads the meaningful unit is the task, not the request. A task consuming ~150K input / 8K output tokens (80% cached, 1.75× reasoning-token multiplier) costs $0.0257 on DeepSeek V4 Pro, $0.097 on Haiku 4.5, $0.254 on Sonnet 5, $0.63 on GPT-5.6 Sol, and $1.27 on Claude Fable 5.
Heavy users run many tasks per day. One intensive coding-agent session on a flagship model burns $2-5, so five sessions a day across 20 working days is $200-500 per user per month. That is not a rounding error on a $20 or even $200 plan. It is the whole unit-economics question.
Usage distributions are long-tailed, so the average user is a misleading design target. Concretely, on a $20 seat with Sonnet 5, a medium user contributes $18.13 of monthly margin. A single heavy agent user burning $200-500 erases the contribution of 11 to 27 of them. Even within pure chat, a 1,000-request user on Fable 5 costs $31.13, underwater on the same $20 seat.
Price, caps, and plan boundaries should be set against the 95th percentile of usage. If your P95 user is unprofitable, no volume of average users fixes it. Growth just scales the leak.
Base per-token prices refresh weekly from the OpenRouter model feed (last refresh 2026-07-11). Everything the feed can’t see, such as caching mechanics, batch discounts, long-context tiers, and announced price changes, is verified by hand against each provider’s official API pricing page.