AI Cost Per User Calculator for LLM Unit Economics

Calculate what each user costs you in LLM API spend. Light, medium, and heavy usage tiers, agent workloads, and why P95 usage matters more than the average.

Model$ / 1M in · outCost / requestCost / 1K requestsCost / month
GPT-5.6 Sol OpenAI$5.00 · $30.00$0.017750$17.750$1775.00
GPT-5.6 Terra OpenAI$2.50 · $15.00$0.008875$8.875$887.50
GPT-5.6 Luna OpenAI$1.00 · $6.00$0.003550$3.550$355.00
GPT-5.4 mini OpenAI$0.75 · $4.50$0.002662$2.663$266.25
GPT-5.4 nano OpenAI$0.20 · $1.25$0.000735$0.735$73.50
Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3×$10.00 · $50.00$0.031125$31.125$3112.50
Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3×$5.00 · $25.00$0.015562$15.562$1556.25
Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled$2.00 · $10.00$0.009338$9.338$933.75
Claude Haiku 4.5 Anthropic (Claude)$1.00 · $5.00$0.003112$3.112$311.25
Gemini 3.1 Pro (Preview) Google (Gemini API)$2.00 · $12.00$0.007100$7.100$710.00
Gemini 3.5 Flash Google (Gemini API)$1.50 · $9.00$0.005325$5.325$532.50
Gemini 2.5 Flash-Lite Google (Gemini API)$0.10 · $0.40$0.000255$0.255$25.50
DeepSeek V4 Pro DeepSeek$0.43 · $0.87$0.000654$0.654$65.43
DeepSeek V4 Flash DeepSeek$0.14 · $0.28$0.000211$0.211$21.14
Mistral Medium 3.5 Mistral$1.50 · $7.50$0.005250$5.250$525.00
Mistral Large 3 Mistral$0.50 · $1.50$0.001025$1.025$102.50
Mistral Small 4 Mistral$0.15 · $0.60$0.000382$0.382$38.25
Grok 4.5 xAI (Grok)$2.00 · $6.00$0.004250$4.250$425.00
Grok 4.3 xAI (Grok)$1.25 · $2.50$0.001975$1.975$197.50
Llama 4 Scout Meta Llama (hosted)$0.10 · $0.30$0.000250$0.250$25.00
Llama 4 Maverick Meta Llama (hosted)$0.15 · $0.60$0.000450$0.450$45.00
Amazon Nova Pro Amazon (Nova on Bedrock)$0.80 · $3.20$0.002400$2.400$240.00
Amazon Nova Lite Amazon (Nova on Bedrock)$0.06 · $0.24$0.000180$0.180$18.00
Amazon Nova Micro Amazon (Nova on Bedrock)$0.04 · $0.14$0.000105$0.105$10.50

The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.

What should you charge?

API cost / user / month $2.80
Gross margin at your price 86.0%
Break-even price $2.80
Price for 70% margin $9.34
Price for 80% margin $14.01
Margin if usage doubles 72.0%
Margin at 5× usage 30.0%

Gross margin counts API cost only. Hosting, support, and payment fees come out of the same price. Healthy AI products keep API costs to 15–30% of revenue, and a common rule of thumb is charging 3.5–5× your token cost. The usage-growth rows show what happens to margin when your users get more active without paying more.

Charging for usage means metering it. Kelviq bundles usage metering, entitlements, and merchant-of-record billing for AI products into one transaction fee.

Cost per user is one multiplication and one guess

The formula is short. Cost per user = cost per request × requests per user per month. The first factor is knowable to four decimal places, and the calculator above computes it live for 24 models. The second factor is the one founders get wrong, because before launch nobody knows how often users will actually hit the product, and after launch the average hides a long tail.

Everything below uses the standard chat anchor (1,000 input / 500 output tokens per request, 50% cached), priced July 11, 2026.

Light, medium, and heavy users

A practical planning split is 100 / 300 / 1,000 requests per user per month.

ModelLight (100/mo)Medium (300/mo)Heavy (1,000/mo)
Claude Fable 5$3.11$9.34$31.13
GPT-5.6 Sol$1.78$5.33$17.75
Claude Sonnet 5$0.62$1.87$6.23
GPT-5.6 Luna$0.36$1.07$3.55
Claude Haiku 4.5$0.31$0.93$3.11
GPT-5.4 mini$0.27$0.80$2.66
DeepSeek V4 Pro$0.07$0.20$0.65

The benchmark that turns these into decisions is that healthy AI products keep API costs to 15-30% of revenue. On a $20 seat, that is a $3-6 monthly cost budget per user. Sonnet 5 sits comfortably inside it at medium usage ($1.87), breaches it with heavy users ($6.23), and Fable 5 blows through it for anyone above light usage.

Agent products change the unit entirely

For agent workloads the meaningful unit is the task, not the request. A task consuming ~150K input / 8K output tokens (80% cached, 1.75× reasoning-token multiplier) costs $0.0257 on DeepSeek V4 Pro, $0.097 on Haiku 4.5, $0.254 on Sonnet 5, $0.63 on GPT-5.6 Sol, and $1.27 on Claude Fable 5.

Heavy users run many tasks per day. One intensive coding-agent session on a flagship model burns $2-5, so five sessions a day across 20 working days is $200-500 per user per month. That is not a rounding error on a $20 or even $200 plan. It is the whole unit-economics question.

Why P95 matters more than the average

Usage distributions are long-tailed, so the average user is a misleading design target. Concretely, on a $20 seat with Sonnet 5, a medium user contributes $18.13 of monthly margin. A single heavy agent user burning $200-500 erases the contribution of 11 to 27 of them. Even within pure chat, a 1,000-request user on Fable 5 costs $31.13, underwater on the same $20 seat.

Price, caps, and plan boundaries should be set against the 95th percentile of usage. If your P95 user is unprofitable, no volume of average users fixes it. Growth just scales the leak.

The levers that actually move cost per user

  • Caching. The anchors here assume 50% cache hits for chat and 80% for agents, so a workload with no cache reuse costs materially more per request. Structuring prompts for cache stability is the cheapest optimization available.
  • Model routing. Classify requests and send the simple ones to Haiku 4.5, GPT-5.4 mini, or DeepSeek V4 Pro. The spread between the cheapest and most expensive model in the table above is roughly 45× at identical usage.
  • Usage caps. A cap converts the open-ended P95 problem into a bounded one, and pairs naturally with credit or hybrid pricing, which the usage-based pricing calculator covers.

Price sources & freshness

Base per-token prices refresh weekly from the OpenRouter model feed (last refresh 2026-07-11). Everything the feed can’t see, such as caching mechanics, batch discounts, long-context tiers, and announced price changes, is verified by hand against each provider’s official API pricing page.

  • OpenAI official pricing, details last verified 2026-07-11 (The GPT-5.6 family (Sol/Terra/Luna) launched Jul 9, 2026. Its three-tier naming replaces the mini/nano ladder at the top end.)
  • Anthropic (Claude) official pricing, details last verified 2026-07-11 (Opus 4.7+, Sonnet 5, and Fable 5 use a new tokenizer producing up to roughly 30% more tokens for the same text, so cross-provider per-token comparisons understate Claude cost slightly.)
  • Google (Gemini API) official pricing, details last verified 2026-07-11 (Gemini 3.1 Pro doubles input pricing above 200K input tokens, the only two-tier context pricing in this comparison.)
  • DeepSeek official pricing, details last verified 2026-07-11 (Prices below are DeepSeek's first-party API. OpenRouter routes some DeepSeek models through discounted hosts at different prices.)
  • Mistral official pricing, details last verified 2026-07-11 (Verify against the API pricing page, since Mistral's marketing pages have shown stale prices for retired models.)
  • xAI (Grok) official pricing, details last verified 2026-07-11
  • Meta Llama (hosted) official pricing, details last verified 2026-07-11 (These are open-weight models, so the same model costs different amounts on different hosts such as Groq, DeepInfra, and Bedrock. Prices shown are OpenRouter's routed rate, indicative rather than contractual.)
  • Amazon (Nova on Bedrock) official pricing, details last verified 2026-07-11 (Bedrock also hosts Anthropic, Mistral, and DeepSeek models, with Claude at first-party prices and DeepSeek at different ones. Only Amazon's own Nova line is listed here.)

Frequently asked questions

Multiply cost per request by requests per user per month. A chat request of 1,000 input and 500 output tokens with 50% caching costs $0.006225 on Claude Sonnet 5, so a user making 300 such requests costs $1.87 per month. The hard part is estimating requests per user, which most founders guess badly.

For a chat product it ranges from under $0.10 to over $30 depending on model and usage. At 300 requests per month, GPT-5.6 Luna costs $1.07 per user, Claude Sonnet 5 costs $1.87, and Claude Fable 5 costs $9.34. Agent products are far higher because each task can cost $0.10 to $1.27 on its own.

A useful planning split is 100 requests per month for light users, 300 for medium, and 1,000 for heavy chat users. Real distributions are heavily skewed, so median users sit near the light tier while a small share of heavy users dominates total spend. Instrument actual usage as early as possible rather than relying on averages.

Usage is long-tailed, so pricing set against the average leaves you exposed to the users at the 95th percentile. One heavy coding-agent user burning $200-500 per month on a flagship model can erase the margin contribution of 11 to 27 average users on a $20 plan. Price and cap against P95, not the mean.

An agent task of about 150K input and 8K output tokens (80% cached, with reasoning overhead) costs $0.0257 on DeepSeek V4 Pro, $0.097 on Claude Haiku 4.5, $0.254 on Claude Sonnet 5, $0.63 on GPT-5.6 Sol, and $1.27 on Claude Fable 5. One intensive coding-agent session on a flagship model burns $2-5, so a heavy user running 5 sessions a day for 20 working days costs roughly $200-500 per month.

The three practical levers are prompt caching, model routing, and usage caps. The figures on this page already assume a 50% cache hit rate for chat and 80% for agent workloads, so uncached workloads cost meaningfully more. Routing simple requests to a cheaper model, say Haiku 4.5 or DeepSeek V4 Pro instead of a flagship, often cuts cost per user by 70-90% for those requests.