Costs & ROI8 min read

How to Reduce LLM API Costs: 12 Practical Tactics

Twelve practical ways to reduce LLM API costs without hurting quality: model routing, prompt caching, shorter context, output limits, batching and monitoring.

To reduce LLM API costs, start with the three biggest levers: use the smallest model that does the job well, send less input per request, and limit output length. Then add structural savings such as prompt caching, batch processing and caching full answers, and keep it all under control with per-feature cost monitoring and spending limits. The twelve tactics below are ordered roughly by typical impact. None of them requires sacrificing quality if you test each change against a fixed set of real cases.

If you are not yet sure how your bill is composed, read LLM API pricing explained first. To see the effect of each tactic on your numbers, use the LLM API cost calculator and change one input at a time.

Before you optimise: measure

Cost optimisation without data is guesswork. Log these fields for every request:

Field Why it matters
Feature or workflow step Shows which part of the product drives cost
Model Shows whether expensive models handle cheap tasks
Input tokens Reveals oversized prompts and context
Output tokens Reveals verbose answers
Cached tokens (if applicable) Shows whether caching works
Retries Reveals failure loops
Latency Often improves alongside cost

In most applications, a small number of steps cause most of the spend. Optimise those first.

1. Route each task to the smallest model that works

Model choice usually has the largest single effect, because prices between small and large models can differ by an order of magnitude or more. Many tasks do not need the most capable model:

  • Small models: classification, routing, extraction of clear fields, short rewrites, simple summaries.
  • Mid-size models: most drafting, standard customer replies, structured analysis.
  • Large models: complex reasoning, nuanced writing, difficult edge cases.

A common pattern is a router: a cheap model or simple rules decide which model handles each request. Another is escalation: try the small model first and pass the case to a larger one only if a check fails. How to choose an LLM describes how to evaluate candidates on your own tasks.

2. Trim the context you send

Input is often the larger part of the bill because system prompts, history and documents are resent with every request. Ways to cut it:

  • Send only the relevant sections of documents, not entire files. Retrieval that selects the best passages is usually cheaper and more accurate than sending everything.
  • Remove boilerplate, navigation text, signatures and duplicated content from inserted material.
  • Prefer compact formats. A clean list of fields uses fewer tokens than a verbose HTML table.
  • Tighten the system prompt. Remove repeated or contradictory instructions; keep examples that earn their place.

3. Manage conversation history

In chat applications, each turn usually resends the whole history, so cost per message grows over a conversation. Options:

  • Keep only the last few turns verbatim.
  • Replace older turns with a short running summary.
  • Store facts (customer name, order number) as structured state rather than relying on the full transcript.

4. Use prompt caching where it applies

Many providers offer prompt caching: if consecutive requests start with an identical prefix, that part can be billed at a reduced rate. To benefit:

  • Put stable content first: system prompt, tool definitions, reference documents.
  • Put variable content last: the user's message and request-specific data.
  • Keep the prefix byte-for-byte identical. A timestamp or user name at the top breaks the cache.
  • Check the minimum length, how long the cache lives and whether writing to it costs extra.

Monitor the cached token counts to confirm it is working.

5. Limit and shape the output

Output tokens are usually the most expensive per token. Simple instructions help:

  • State a length: "Answer in at most three sentences" or "maximum 150 words".
  • Ask for the format you need and nothing else: "Return only the JSON object, without explanations".
  • Set a maximum output token limit as a safety net.
  • Avoid asking the model to repeat the input back.

Shorter outputs are also faster to read and review, which saves people's time as well.

6. Control reasoning effort

Some models can spend extra tokens on internal reasoning before answering. That improves results on hard problems but is wasted on simple ones. If your provider offers settings for reasoning effort or a thinking budget, use low settings for routine tasks and higher ones only where testing shows a real quality gain.

7. Batch non-urgent work

If results are not needed immediately, such as overnight report generation, bulk classification or data enrichment, check whether your provider offers a batch interface. Batch processing is commonly offered at a discount compared with real-time requests, in exchange for slower turnaround.

8. Cache complete answers

If users often ask the same questions, store the answer and reuse it instead of calling the model again. Exact-match caching is simple. Semantic caching, which reuses answers for similar questions, saves more but needs care so that subtly different questions do not get the wrong answer. Always set an expiry so cached answers do not outlive the underlying facts.

9. Avoid unnecessary calls

Not every step needs a language model:

  • Use regular code for deterministic tasks such as date formatting, calculations or validation.
  • Use simple rules or keyword filters before the model, for example to discard spam.
  • Combine steps when it does not hurt quality. One call that classifies and extracts can replace two.

Our comparison of AI agents and workflow automation explains why fixed workflows with a few targeted model calls are often cheaper and more predictable than open-ended agents.

10. Reduce retries and failures

Every failed attempt costs tokens. Common causes and fixes:

  • Malformed structured output: use the provider's structured output or schema features where available, and give a clear example.
  • Vague prompts leading to rejected drafts: improve the prompt; see the prompt engineering guide.
  • Agent loops: set a maximum number of steps and a budget per task.
  • Timeouts with automatic retries: make sure retries are not multiplying requests unexpectedly.

11. Consider fine-tuning or smaller specialised models for high volume

For a narrow, high-volume task, a smaller model fine-tuned on good examples can sometimes match a larger model at lower cost per request, and with shorter prompts because instructions no longer need to be repeated. Fine-tuning has its own costs for training, evaluation and maintenance, so it pays off only at sufficient volume. RAG vs. fine-tuning covers when it makes sense.

12. Set budgets, alerts and reviews

Optimisation is not a one-off project. Put guardrails in place:

  • Spending limits and alerts in the provider console.
  • Per-user or per-tenant quotas in your application.
  • A monthly review of cost per feature and cost per successful outcome.
  • A check of the price list whenever you plan a change. Providers regularly release new models and adjust prices.

Worked example: combining tactics

Assume a support assistant with these example figures: 3,000 input tokens and 400 output tokens per request, 20,000 requests per month, using a mid-size model at example prices of 1.00 per million input tokens and 4.00 per million output tokens.

  • Baseline: (3,000 × 1.00 + 400 × 4.00) ÷ 1,000,000 = 0.0046 per request, about 92 per month.

Now apply three tactics, with assumed effects:

  1. Trim retrieved context from 1,800 to 900 tokens: input falls to 2,100 tokens.
  2. Cache the 800-token system prompt at an example cached price of 0.25: those tokens cost a quarter.
  3. Limit answers: output falls to 250 tokens.
  • New input cost: (1,300 × 1.00 + 800 × 0.25) ÷ 1,000,000 = 0.0015
  • New output cost: 250 × 4.00 ÷ 1,000,000 = 0.001
  • New total: 0.0025 per request, about 50 per month.

That is a reduction of nearly half without changing the model. Routing simple questions to a smaller model would reduce it further. Your numbers will differ; the method is what transfers.

What not to do

  • Cut quality to save pennies. If a cheaper setup produces more errors, the human correction time can cost far more than the tokens saved. Compare total cost including review, for example with the AI ROI calculator.
  • Optimise without a test set. Every change should be checked against the same real cases.
  • Strip out safety instructions. Rules about data handling and accuracy belong in the prompt even if they cost tokens.
  • Rely on prices from memory. Always check the current price list.

Summary checklist

  • Log tokens and cost per feature
  • Use the smallest model that passes your tests
  • Send only relevant context
  • Summarise long histories
  • Put stable content first and enable caching
  • Limit output length and format
  • Use low reasoning effort for routine tasks
  • Batch non-urgent jobs
  • Cache repeated answers
  • Replace deterministic steps with code
  • Cap retries and agent steps
  • Set budgets and review monthly

To estimate the effect of these changes, enter before-and-after values in the LLM API cost calculator and compare the scenarios side by side.

FAQ

What is the fastest way to reduce LLM API costs?

Check which model you use for each task. Routing simple tasks to a smaller, cheaper model and reserving large models for hard cases often has the biggest effect. Trimming unnecessary context is a close second.

Does prompt caching really save money?

It can, when many requests share the same long prefix such as a system prompt or reference document. Cached input is billed at a reduced rate, but the prefix must be identical and placed at the start of the prompt.

Will cheaper models reduce quality?

Sometimes, which is why you should test them on your own cases. For classification, extraction and short drafting, smaller models are often good enough. Measure quality before and after any switch.

How do I find out where my LLM costs come from?

Log input and output tokens per request together with the feature, model and user action. Grouping costs by feature usually reveals that a few steps cause most of the spend.

Related articles

Costs & ROI8 min read

How to Calculate the ROI of AI Automation

A step-by-step method to calculate AI automation ROI: time saved, review effort, running and setup costs and payback, with a worked example and common traps.

← Back to the blog