AI tokens: What they are and what they cost
What if AI didn't increase your bill? Learn how tokens work and how to control costs.
What are AI tokens?
Every time you interact with an AI model, the system isn't reading words. It's reading tokens.
An AI token is the basic unit of text that a large language model (LLM) processes. Tokens aren't always whole words. They're chunks of text: sometimes a full word, sometimes a syllable, sometimes a single punctuation mark. The word "software" might be one token. "Tokenization" might be three.
Before any AI reasoning happens, your input is broken into chunks and converted into numbers. The model thinks in tokens. It responds in tokens. And in most AI pricing models, you pay in tokens.
Why tokens impact your business
Token counts measure every unit of text an AI model reads or writes. Each interaction is billed at tokens consumed × price per token. In agentic systems, agents re-read full context at every reasoning step, and costs compound fast. This is what Pega calls the AI token tax: the hidden cost that makes agentic AI hard to budget and scale. Controlling token counts means controlling your AI spend.
How tokenization works
Tokenization converts raw text into numerical sequences that AI models can process, using a method called Byte-pair encoding (BPE):
- Text is broken into characters. Every letter, space, and punctuation mark starts as a candidate unit.
- Frequent pairs are merged. Common adjacent characters combine into single tokens ("t" + "h" → "th" → "the").
- A fixed vocabulary is built. Roughly 100,000 unique tokens cover common words, sub-words, code, and multiple languages.
- Text becomes numbers. Each token is assigned a numerical identifier. The model processes these – not the original text.
The practical result: Common words are single tokens. Rare or technical terms fragment into several – and cost more to process.
How much money are you wasting on AI re-reasoning?
Token limits: Context windows explained
Every AI model has a context window – the maximum number of tokens it can process in a single interaction. Think of it as the model's working memory. Anything outside the window is invisible to the model.
As of 2026, context windows have expanded dramatically:
- Standard enterprise models support between 200,000 and 1,000,000 tokens per interaction.
- Extended-context models offer windows of 2 million to 10 million tokens – enough to process entire codebases or hours of recorded calls in a single prompt.
Larger windows are powerful – and expensive. Filling them with redundant instructions or unfiltered data wastes budget and degrades output quality. Precision matters more than capacity.
Input vs. output tokens
Not all tokens are priced equally. Understanding the difference between input and output tokens is essential for accurate cost forecasting.
Tokenization in practice
The AI token tax: The hidden cost of agentic AI
In traditional AI usage, token costs are predictable: one prompt, one response, one cost. Agentic AI doesn't work that way. Autonomous agents plan and reason across multiple steps, and at every step they re-read their full context to maintain coherence. Each re-read is a token event. Each pass adds to the bill.
The result: LLM-orchestrated workflows cost 5 to 20X more per run than targeted execution. And the gap compounds with every additional workflow step, as context windows grow and token usage accelerates.
The AI token tax compounds in three ways:
- Depth: More reasoning steps = more output tokens.
- Width: More agents in parallel = more simultaneous token events.
- Recurrence: Agents rereading shared context at every step = more input tokens per step.
Understanding the token tax is the first step to building AI that scales without surprise costs.
What organizations can do to manage token costs
Predictable AI: No token fees. No surprises.
The alternative to the token tax isn't fewer AI features. It's smarter AI architecture.
Pega's Predictable AI model reframes the token question entirely. Rather than asking "How many tokens did this cost?" it asks, "What outcome did this achieve?" Pricing is anchored to the completed unit of work – the case – not the computational path taken to complete it.
This means:
- No token-based billing surprises. Costs are fixed for the duration of the contract.
- No usage caps. Teams can use as much AI as the work requires without watching a meter.
- Full governance. Because runtime logic remains deterministic, workflows are auditable, compliant, and traceable – even when AI does the heavy lifting at design time.
The Predictable AI model is built for enterprises that need AI to be both powerful and governable – not a trade-off between capability and cost control.