Costs & ROI8 min read

LLM API Pricing Explained: How Token Costs Add Up

How LLM API pricing works: input and output tokens, context, caching, batch discounts and hidden multipliers, with a worked monthly cost example.

LLM API pricing is usage-based: you pay for the tokens you send to the model (input) and the tokens it generates (output), each at its own price, usually quoted per million tokens. The cost of one request is input tokens × input price plus output tokens × output price. Your monthly bill is that figure multiplied by the number of requests, plus any extras such as tool calls or storage. The rest of this guide explains what drives each part of the formula and where estimates usually go wrong.

This is the hub article for our cluster on AI costs. If you want to calculate as you read, open the LLM API cost calculator in a second tab.

The basic pricing model

Language models do not read words; they read tokens, which are chunks of text such as whole short words, parts of longer words, punctuation or spaces. If tokens are new to you, read what tokens are in LLMs first. For pricing, three facts matter:

  1. Every request is billed by tokens. Both directions count.
  2. Input and output have different prices. Output is usually several times more expensive per token.
  3. Prices are quoted per million tokens. A price of 2.00 per million input tokens means 0.000002 per token.

So the core formula is:

Cost per request = (input tokens × input price + output tokens × output price) ÷ 1,000,000

Prices differ widely between providers and between models of the same provider. Smaller, faster models can cost a fraction of the largest ones. Because these prices change frequently, check the provider's official pricing page before you budget; any number you read in an article, including the examples here, should be treated as illustrative.

What counts as input

Input is everything you send with a request, not only the user's question. In a typical business application, that includes:

  • System prompt: the standing instructions that define the assistant's role, rules and format.
  • Conversation history: in a chat, previous messages are usually resent each turn so the model has context.
  • Retrieved context: documents, knowledge base passages or database records inserted into the prompt, for example in a retrieval-augmented generation (RAG) setup.
  • Tool definitions: if the model can call functions, their descriptions are part of the input.
  • Images or files: many models convert these into tokens as well, often at a significant count per image or page.
  • The actual user message.

This is why input is often the larger share of costs even though its per-token price is lower. A short user question can travel with thousands of tokens of instructions and context.

What counts as output

Output is the text the model generates in its answer. Two things make it easy to underestimate:

  • Verbose answers. Without guidance on length, models tend to be generous. A clear format instruction can cut output substantially.
  • Reasoning tokens. Some models "think" before answering and generate internal reasoning tokens. Depending on the provider, these are billed as output even if you do not see them in the response. Check how your chosen model reports and bills them.

Pricing features that change the calculation

Beyond the basic rates, most providers offer pricing mechanisms that can lower or raise your costs. Names and conditions vary, so read the documentation of your provider.

Feature What it does Effect on cost
Prompt caching Reuses an identical prompt prefix across requests Cached input is billed at a reduced rate; writing to the cache may cost extra
Batch processing Requests are processed asynchronously, not in real time Often discounted compared with standard requests
Context length tiers Some models charge more above a certain prompt length Very long prompts can cost more per token
Tool and search calls Built-in web search, code execution or file search Often billed per call or per use on top of tokens
Fine-tuned models Models trained on your data Training cost plus usually different inference prices
Priority or reserved capacity Guaranteed throughput Higher or committed spending

Caching deserves special attention. If many requests share a long system prompt or the same reference document, caching can reduce the input cost of that shared part considerably. The calculator lets you enter the share of input served from cache and a separate cached price. Our article on reducing LLM API costs covers how to structure prompts so caching actually applies.

A worked example

Let's estimate a customer support assistant. All figures are assumptions for illustration, and the prices are example values, not those of any real provider.

Usage assumptions

  • System prompt and rules: 800 tokens
  • Retrieved help-centre passages: 1,200 tokens
  • Conversation history (average): 600 tokens
  • User message: 100 tokens
  • Answer: 300 tokens
  • Requests: 1,000 per day, 30 days per month

Per request

  • Input: 800 + 1,200 + 600 + 100 = 2,700 tokens
  • Output: 300 tokens

Example prices: 1.00 per million input tokens, 4.00 per million output tokens.

  • Input cost: 2,700 × 1.00 ÷ 1,000,000 = 0.0027
  • Output cost: 300 × 4.00 ÷ 1,000,000 = 0.0012
  • Cost per request: 0.0039

Per month: 0.0039 × 1,000 × 30 = 117

Notice that input accounts for about 70% of the cost even though output tokens are four times more expensive. Shortening the retrieved context or caching the system prompt would have more effect here than shortening the answers.

Now change one assumption. If the team switches to a larger model with example prices of 5.00 and 20.00, the same usage costs 585 per month. If they use a smaller model at 0.20 and 0.80, it costs about 23. That spread is typical: model choice is often the single biggest cost lever, which is why choosing the right LLM is partly a pricing decision.

Subscriptions versus API pricing

Many people first meet AI through chat subscriptions with a flat monthly fee per user. API access is different:

Chat subscription API
Billing Fixed fee per user per month Pay per token used
Best for Staff using an assistant directly Applications, automations, integrations
Cost predictability High Depends on usage, needs monitoring
Control over prompts and data flow Limited Full, within the provider's terms
Usage limits Fair-use limits set by the provider Rate limits and spending limits you can configure

A business often uses both: subscriptions for knowledge workers and the API for automated processes. Do not compare them only on price. A subscription does not let you build a workflow, and an API does not give staff a ready-made interface.

Hidden multipliers to plan for

The formula is simple; real usage is not. These factors regularly push actual costs above the first estimate:

  • Retries and errors. Failed or rejected outputs that are generated again cost tokens twice.
  • Agent loops. An AI agent may call the model many times to complete one task: planning, calling tools, reading results and checking its work. One user request can mean ten or more model calls. Read AI agents vs. workflow automation for how this affects design choices.
  • Growing conversation history. In long chats, every turn resends the history, so cost per message grows over the conversation.
  • Multiple steps per process. A pipeline that classifies, then extracts, then drafts, then checks makes four calls per document.
  • Development and testing. Prompt experiments and evaluation runs cost money before the first real user arrives.
  • Peak growth. Successful features get used more. Budget for growth, not just the pilot.

A practical rule is to measure real token counts from a sample of test requests rather than estimating from word counts, and then add a margin for the items above.

How to estimate your own costs

  1. Collect sample requests. Take ten to twenty realistic examples of what the application will send and receive.
  2. Measure tokens. Use your provider's token counting tool or the usage figures returned by the API. For a quick first estimate, the token estimator in the cost calculator gives a rough figure from sample text.
  3. Calculate per request. Average input and output tokens, multiplied by the current prices.
  4. Multiply by volume. Requests per day × days per month. Consider calls per user action, not just user actions.
  5. Compare scenarios. Try at least two models. The calculator shows three side by side.
  6. Add a margin. For retries, growth and testing.
  7. Set spending limits. Most providers let you configure budget alerts or hard limits. Use them from day one.

Common mistakes

  • Counting only the user message as input. The system prompt and context are usually larger.
  • Using last year's prices. Prices change often, frequently downward for older models, sometimes with new pricing structures. Always check the current price list.
  • Ignoring output length. One sentence in the prompt about the desired length can save more than switching models.
  • Forgetting non-token charges. Tool calls, storage, fine-tuning and data transfer may be billed separately.
  • Optimising cost before quality. The cheapest model is only cheaper if its answers are good enough. A model that needs more retries or more human correction may cost more overall.

Next steps

To turn your estimate into a business decision, combine the cost side with the value side: the AI ROI calculator compares running costs with the time saved. If the estimate is higher than you hoped, the twelve tactics in how to reduce LLM API costs are the place to continue.

FAQ

How is LLM API usage billed?

Most providers bill per token, with separate prices for input tokens (what you send) and output tokens (what the model generates), usually quoted per million tokens. Some features such as caching, batch processing or tool use have their own rates.

Why is the output price higher than the input price?

Generating tokens one by one takes more compute than processing the input, so providers typically charge more for output. Long answers can therefore dominate the bill.

What makes an LLM application more expensive than expected?

Usually the input side: long system prompts, conversation history and retrieved documents are resent with every request. Retries, agent loops and hidden reasoning tokens can multiply costs as well.

How do I estimate my monthly LLM API cost?

Multiply average input and output tokens per request by their prices, then by the number of requests per month. Use real sample requests to measure token counts and add a safety margin for growth and retries.

Related articles

Costs & ROI8 min read

How to Calculate the ROI of AI Automation

A step-by-step method to calculate AI automation ROI: time saved, review effort, running and setup costs and payback, with a worked example and common traps.

← Back to the blog