Work backwards from LLM API costs to what your AI app should charge. Live margins for 24 models, break-even prices, and 70-80% gross margin targets.
| Model | $ / 1M in · out | Cost / request | Cost / 1K requests | Cost / month |
|---|---|---|---|---|
| GPT-5.6 Sol OpenAI | $5.00 · $30.00 | $0.017750 | $17.750 | $1775.00 |
| GPT-5.6 Terra OpenAI | $2.50 · $15.00 | $0.008875 | $8.875 | $887.50 |
| GPT-5.6 Luna OpenAI | $1.00 · $6.00 | $0.003550 | $3.550 | $355.00 |
| GPT-5.4 mini OpenAI | $0.75 · $4.50 | $0.002662 | $2.663 | $266.25 |
| GPT-5.4 nano OpenAI | $0.20 · $1.25 | $0.000735 | $0.735 | $73.50 |
| Claude Fable 5 Anthropic (Claude) · tokenizer ~1.3× | $10.00 · $50.00 | $0.031125 | $31.125 | $3112.50 |
| Claude Opus 4.8 Anthropic (Claude) · tokenizer ~1.3× | $5.00 · $25.00 | $0.015562 | $15.562 | $1556.25 |
| Claude Sonnet 5 Anthropic (Claude) · tokenizer ~1.3× · price change scheduled | $2.00 · $10.00 | $0.009338 | $9.338 | $933.75 |
| Claude Haiku 4.5 Anthropic (Claude) | $1.00 · $5.00 | $0.003112 | $3.112 | $311.25 |
| Gemini 3.1 Pro (Preview) Google (Gemini API) | $2.00 · $12.00 | $0.007100 | $7.100 | $710.00 |
| Gemini 3.5 Flash Google (Gemini API) | $1.50 · $9.00 | $0.005325 | $5.325 | $532.50 |
| Gemini 2.5 Flash-Lite Google (Gemini API) | $0.10 · $0.40 | $0.000255 | $0.255 | $25.50 |
| DeepSeek V4 Pro DeepSeek | $0.43 · $0.87 | $0.000654 | $0.654 | $65.43 |
| DeepSeek V4 Flash DeepSeek | $0.14 · $0.28 | $0.000211 | $0.211 | $21.14 |
| Mistral Medium 3.5 Mistral | $1.50 · $7.50 | $0.005250 | $5.250 | $525.00 |
| Mistral Large 3 Mistral | $0.50 · $1.50 | $0.001025 | $1.025 | $102.50 |
| Mistral Small 4 Mistral | $0.15 · $0.60 | $0.000382 | $0.382 | $38.25 |
| Grok 4.5 xAI (Grok) | $2.00 · $6.00 | $0.004250 | $4.250 | $425.00 |
| Grok 4.3 xAI (Grok) | $1.25 · $2.50 | $0.001975 | $1.975 | $197.50 |
| Llama 4 Scout Meta Llama (hosted) | $0.10 · $0.30 | $0.000250 | $0.250 | $25.00 |
| Llama 4 Maverick Meta Llama (hosted) | $0.15 · $0.60 | $0.000450 | $0.450 | $45.00 |
| Amazon Nova Pro Amazon (Nova on Bedrock) | $0.80 · $3.20 | $0.002400 | $2.400 | $240.00 |
| Amazon Nova Lite Amazon (Nova on Bedrock) | $0.06 · $0.24 | $0.000180 | $0.180 | $18.00 |
| Amazon Nova Micro Amazon (Nova on Bedrock) | $0.04 · $0.14 | $0.000105 | $0.105 | $10.50 |
The table starts with a chat assistant workload and lists the best-known models first. Adjust the inputs above and every cost recomputes live. Base prices refresh weekly from a live feed (last update 2026-07-11), while caching mechanics, batch discounts, and announced price changes are verified by hand against official pricing pages listed in the sources. The reasoning multiplier applies only to models that bill thinking tokens as output.
Gross margin counts API cost only. Hosting, support, and payment fees come out of the same price. Healthy AI products keep API costs to 15–30% of revenue, and a common rule of thumb is charging 3.5–5× your token cost. The usage-growth rows show what happens to margin when your users get more active without paying more.
Charging for usage means metering it. Kelviq bundles usage metering, entitlements, and merchant-of-record billing for AI products into one transaction fee.
Cost per request is an input, not a decision. The decision is gross margin, meaning (price − API cost per user) ÷ price. Healthy AI products keep API costs to 15-30% of revenue. Typical AI wrappers today run 25-60% gross margins, against 80-90% for classic SaaS, and that gap is what pricing has to close, because the model bill scales with every active user in a way traditional hosting never did.
Most calculators stop at “your workload costs $X.” This one inverts the math. Given a cost per user, it returns the break-even price, the price for a 70% or 80% margin, and what happens to your margin if usage doubles or quintuples.
The inversion is one line. Price = cost per user ÷ (1 − target margin).
Take Claude Sonnet 5 at $0.006225 per chat request (1,000 input / 500 output tokens, 50% cached). At 300 requests per user per month, the numbers fall out like this.
That is where the 3.5-5× rule of thumb comes from, since charging 3.5-5 times token cost yields 70-80% gross margin. On the cheap end, GPT-5.6 Luna costs $1.07/user at the same usage, a 94.7% margin at $20. One caveat is baked into the numbers. Sonnet 5’s figure is intro pricing that rises about 50% on September 1, 2026, which pushes cost per user to roughly $2.80 and trims the $20-seat margin to 86%.
A flat seat is the simplest sell, but margin then floats with each user’s consumption. Credit packs, the most common AI wrapper pricing model, cap your exposure. The user prepays for a fixed amount of usage, and 92% of AI software companies now mix subscription and usage pricing in some form.
Credits also introduce breakage. Cursor Pro charges $20/month for roughly $20 of included usage, approximately 0% margin on usage actually consumed. The margin comes from users who don’t consume their full allowance. Breakage can rescue an otherwise thin sticker margin, but it is fragile. As power users learn to max out allowances, realized margin converges toward the consumed-usage margin.
The margin panel above shows your margin at 2× and 5× current usage for a reason. At a $20 seat on Sonnet 5, 300 requests/month is a 90.7% margin, 600 requests is 81.3%, and 1,500 requests is 53.3%. On Claude Fable 5 the same seat starts at 53.3% margin, drops to 6.6% at 2× usage, and is deeply underwater at 5×. Flat seats with growing per-user usage is the standard way AI products quietly lose their margins, because usage almost never shrinks after you ship a feature that increases it.
When margin is short of target, you have two levers, and they are not symmetric.
Switch models when quality permits. At a $6 seat with Sonnet 5, margin is 68.9%. Routing to Haiku 4.5 ($0.93/user) takes it to 84.4%, GPT-5.6 Luna ($1.07/user) to 82.2%, and DeepSeek V4 Pro ($0.20/user) to 96.7%. Model routing, sending cheap models the simple requests and flagships the hard ones, captures much of this without a full switch.
Raise price when quality doesn’t. If the product only works on a flagship, the price must reflect it. Fable 5 at $9.34/user needs a seat above $31 for 70% margin. Pricing below that isn’t a growth strategy. It’s a countdown.
Base per-token prices refresh weekly from the OpenRouter model feed (last refresh 2026-07-11). Everything the feed can’t see, such as caching mechanics, batch discounts, long-context tiers, and announced price changes, is verified by hand against each provider’s official API pricing page.