Costs & ROI7 min read

What Are Tokens in LLMs? A Plain-English Explainer

What tokens in LLMs are, how text is split into them, why languages and formats differ, and how tokens drive context limits, speed and API cost, with examples.

A token is the basic unit of text a large language model (LLM) reads and produces. Before a model sees your prompt, the text is split into tokens: common words are often a single token, longer or rarer words are split into several pieces, and spaces, punctuation and numbers are tokens or parts of tokens too. For ordinary English, one token is roughly four characters or three quarters of a word. Tokens matter in practice because they determine what a request costs, how much text fits into the model's context window, and how long an answer takes to generate.

Why models use tokens instead of words or letters

A model needs a fixed vocabulary of units it can represent internally. Using whole words would require an enormous vocabulary and still fail on new words, names and typos. Using single characters would make texts very long sequences and slow everything down.

Tokens are the compromise. A tokenizer learns, from large amounts of text, which character sequences appear often and turns them into vocabulary entries. Frequent words become single tokens. Less frequent words are assembled from smaller pieces. This way, the model can represent any text, including invented words, while keeping common text compact.

Each model family has its own tokenizer, so the same sentence can produce different token counts with different models. That is why exact counts should always come from the tokenizer of the model you actually use.

What tokenization looks like

The exact split depends on the tokenizer, but a typical pattern looks like this (illustrative, not the output of a specific model):

Text Possible tokens Count
The invoice is due The invoice is due 4
Unbelievably Un believ ably 3
2026-10-04 202 6 - 10 - 04 6
Kundenzufriedenheit K unden zuf ried enheit 5

Notice a few patterns:

  • Leading spaces are often part of the token. " invoice" with a space is a different token from "invoice" at the start of a line.
  • Long words split into pieces. That is especially common in languages with compound words.
  • Numbers and dates can be surprisingly expensive. Digits are often grouped in small chunks.

Rules of thumb for estimating tokens

When you need a quick estimate, these rough rules work for planning:

Content Rough estimate
English prose 1 token ≈ 4 characters ≈ 0.75 words
1,000 English words about 1,300–1,400 tokens
One page of English text (around 500 words) about 650–700 tokens
German, French, Spanish often noticeably more tokens than the same content in English
Source code, JSON, tables often more tokens per character than prose, because of symbols and whitespace
Languages with non-Latin scripts can be considerably more, depending on the tokenizer

These are estimates only. For budgets, measure real samples. The LLM API cost calculator includes a quick estimator: paste a typical prompt or answer and it gives a rough token count you can transfer into the cost calculation.

Input tokens and output tokens

Every request has two token counts:

  • Input tokens: everything you send. That includes the system prompt, previous conversation turns, inserted documents, tool definitions and the user's message.
  • Output tokens: everything the model generates in response. Some models also generate internal reasoning tokens before the visible answer, which may be billed as output.

The distinction matters because most providers price them differently, with output usually more expensive per token. The LLM API pricing guide explains how the two add up to your bill, with a worked example.

The context window: tokens as capacity

The context window is the maximum number of tokens a model can consider in one request, input and output combined. If a model has a context window of, say, 128,000 tokens, your prompt plus the answer must fit within that.

What that means in practice:

  • Long documents may not fit. A large contract collection or a long email thread can exceed the window.
  • Chats have a memory limit. In a long conversation, older turns eventually need to be dropped or summarised.
  • A bigger window is not free. Every token you send costs money and time. Sending a whole manual when one section is relevant is wasteful.
  • More context is not always better. Models can overlook details buried in very long inputs. Focused context tends to produce more reliable answers.

When documents are too large or too many, the usual solution is retrieval: search for the relevant passages first and send only those. Our comparison of RAG vs. fine-tuning explains how that works.

Tokens and speed

Models generate output one token at a time. Two practical consequences follow:

  1. Long answers take longer. If users are waiting, asking for concise output improves the experience.
  2. Long inputs add delay too. Processing a very long prompt takes time before the first output token appears, although usually less per token than generation.

If your application feels slow, check token counts before switching providers.

Why the same content can cost more in other languages

Tokenizers are trained on large text collections, and English tends to be well represented. As a result, English text is often tokenized efficiently, while other languages may need more tokens for the same meaning. Compound-heavy languages such as German, or languages with rich inflection, can be noticeably more expensive per sentence.

For a business operating in several languages, this has two implications:

  • Budget per language. A support assistant answering in German may use more tokens per conversation than the same assistant in English.
  • Test with real local content. Do not extrapolate from English samples.

Worked example: from words to cost

Assume an internal assistant that summarises reports. All numbers are example assumptions.

  • Report length: 3,000 English words, so roughly 4,000 tokens.
  • Instructions: 300 tokens.
  • Summary: 250 words, so roughly 330 tokens.
  • Volume: 400 reports per month.

Input per request: about 4,300 tokens. Output: about 330 tokens.

With example prices of 1.00 per million input tokens and 4.00 per million output tokens:

  • Input: 4,300 × 1.00 ÷ 1,000,000 = 0.0043
  • Output: 330 × 4.00 ÷ 1,000,000 = 0.00132
  • Per report: about 0.0056
  • Per month: about 2.25

The cost here is small, but the same arithmetic applied to a high-volume customer chat, or to an agent that makes many calls per task, quickly reaches larger sums. That is why measuring tokens early is worthwhile. For ways to bring the count down, see how to reduce LLM API costs.

How to count tokens precisely

  • Provider tools: most providers offer a tokenizer, a token counting endpoint or a playground that shows token counts.
  • API responses: each response usually includes the number of input and output tokens used. Logging these is the most reliable way to know your real consumption.
  • Open-source tokenizer libraries: for some model families, the tokenizer is available as a library you can run locally.

For planning, rule-of-thumb estimates are fine. For budgets and monitoring, use real counts.

Common misconceptions

  • "A token is a word." Often, but not always. Many words are several tokens, and spaces and punctuation count.
  • "Only my question counts." System prompts, history and documents count as input on every request.
  • "A bigger context window solves everything." It increases capacity, not focus, and raises cost when used.
  • "Token counts are the same across models." Each tokenizer is different.
  • "Shorter prompts are always better." Removing necessary context saves tokens but can cost more in poor results and corrections.

Key takeaways

Tokens are the currency of LLMs: they set the price, the capacity and the speed. Estimate with rules of thumb, measure with real samples, and remember that input includes everything you send, not just the user's words. To see what your own volumes would cost, try the LLM API cost calculator, and read the LLM API pricing guide for the full picture of how providers bill.

FAQ

What is a token in an LLM?

A token is the unit of text a language model reads and writes. It can be a whole word, part of a word, a punctuation mark or a space. Models process and generate text token by token.

How many words is 1,000 tokens?

For typical English text, roughly 700 to 800 words, using the common rule that one token is about three quarters of a word. Other languages, code and numbers often need more tokens for the same content.

Why do tokens matter for businesses?

API usage is billed per token, context windows are measured in tokens, and response time grows with the number of tokens generated. Tokens therefore drive cost, capacity and speed.

How can I count tokens exactly?

Use the tokenizer or token counting endpoint of the specific model you plan to use, or read the usage figures returned with each API response. Rules of thumb are only estimates.

Related articles

Costs & ROI8 min read

How to Calculate the ROI of AI Automation

A step-by-step method to calculate AI automation ROI: time saved, review effort, running and setup costs and payback, with a worked example and common traps.

← Back to the blog