メインコンテンツに飛ぶ

We'd prefer it if you saw us at our best.

Pega.com is not optimized for Internet Explorer. For the optimal experience, please use:

Close Deprecation Notice
このページはあなたの言語(日本語)ではご利用いただけません。元の言語(英語)で以下からお読みいただけます。

AI tokens: What they are and what they cost

What if AI didn't increase your bill? Learn how tokens work and how to control costs.

AI Tokens

What are AI tokens?

Every time you interact with an AI model, the system isn't reading words. It's reading tokens.

An AI token is the basic unit of text that a large language model (LLM) processes. Tokens aren't always whole words. They're chunks of text: sometimes a full word, sometimes a syllable, sometimes a single punctuation mark. The word "software" might be one token. "Tokenization" might be three.

Before any AI reasoning happens, your input is broken into chunks and converted into numbers. The model thinks in tokens. It responds in tokens. And in most AI pricing models, you pay in tokens.

Why tokens impact your business

Token counts measure every unit of text an AI model reads or writes. Each interaction is billed at tokens consumed × price per token. In agentic systems, agents re-read full context at every reasoning step, and costs compound fast. This is what Pega calls the AI token tax: the hidden cost that makes agentic AI hard to budget and scale. Controlling token counts means controlling your AI spend.

AI Tokens Impact

How tokenization works

Tokenization converts raw text into numerical sequences that AI models can process, using a method called Byte-pair encoding (BPE):

  1. Text is broken into characters. Every letter, space, and punctuation mark starts as a candidate unit.
  2. Frequent pairs are merged. Common adjacent characters combine into single tokens ("t" + "h" → "th" → "the").
  3. A fixed vocabulary is built. Roughly 100,000 unique tokens cover common words, sub-words, code, and multiple languages.
  4. Text becomes numbers. Each token is assigned a numerical identifier. The model processes these – not the original text.

The practical result: Common words are single tokens. Rare or technical terms fragment into several – and cost more to process.

AI Tokens How it works

How much money are you wasting on AI re-reasoning?

Calculate your exposure
AI Token Limit

Token limits: Context windows explained

Every AI model has a context window – the maximum number of tokens it can process in a single interaction. Think of it as the model's working memory. Anything outside the window is invisible to the model.

As of 2026, context windows have expanded dramatically:

  • Standard enterprise models support between 200,000 and 1,000,000 tokens per interaction.
  • Extended-context models offer windows of 2 million to 10 million tokens – enough to process entire codebases or hours of recorded calls in a single prompt.

Larger windows are powerful – and expensive. Filling them with redundant instructions or unfiltered data wastes budget and degrades output quality. Precision matters more than capacity.

Input vs. output tokens

Not all tokens are priced equally. Understanding the difference between input and output tokens is essential for accurate cost forecasting.

Input tokens

  • What they are: Text sent to the model
  • Examples: Prompts, instructions, context, history
  • Processing: Parallel (faster, cheaper)
  • Relative cost: Lower

Output tokens

  • What they are: Text generated by the model
  • Examples: Answers, summaries, code, reasoning steps
  • Processing: Sequential (slower, costlier)
  • Relative cost: 3–5X higher

Tokenization in practice

Customer service

An AI support agent processes the customer message, retrieves relevant knowledge articles, reads conversation history, and generates a response. A single interaction: 2,000–5,000 tokens. Token costs become a primary operational expense when they are spread across millions of daily interactions.

Agentic workflows

An AI agent resolving a complex insurance claim reasons across 12 steps – checking policy data, querying systems, drafting responses, verifying compliance. Each step re-reads shared context. Costs stack with every iteration.

Content generation

Long system prompts, injected brand guidelines, and multi-turn revision loops each add input tokens. Output-heavy tasks multiply the output token bill.

The AI token tax: The hidden cost of agentic AI

In traditional AI usage, token costs are predictable: one prompt, one response, one cost. Agentic AI doesn't work that way. Autonomous agents plan and reason across multiple steps, and at every step they re-read their full context to maintain coherence. Each re-read is a token event. Each pass adds to the bill.

The result: LLM-orchestrated workflows cost 5 to 20X more per run than targeted execution. And the gap compounds with every additional workflow step, as context windows grow and token usage accelerates.

The AI token tax compounds in three ways:

  • Depth: More reasoning steps = more output tokens.
  • Width: More agents in parallel = more simultaneous token events.
  • Recurrence: Agents rereading shared context at every step = more input tokens per step.

Understanding the token tax is the first step to building AI that scales without surprise costs.

Ai Token Tax

What organizations can do to manage token costs

Move AI to design time

Build and configure workflows with AI before they go live. Runtime execution stays lean and deterministic – no token overhead per case.

Adopt outcome-based pricing

Per-token billing is unpredictable at scale. Outcome-based models – priced per case or resolution – eliminate the token tax by decoupling cost from token volume.

Use RAG instead of loading everything at once

Think of it like a search engine inside your AI. Instead of dumping an entire database into the model's memory, Retrieval-Augmented Generation (RAG) finds and pulls in only the pages that matter for the question at hand.

Cache common prompts

System instructions and brand guidelines that appear in every interaction can be cached – eliminating the cost of reprocessing them on every call.

Summarize before reasoning

Compress long documents or conversation histories into shorter context before feeding them to the model. Fewer tokens in = lower cost per step.
Predictable AI

Predictable AI: No token fees. No surprises.

The alternative to the token tax isn't fewer AI features. It's smarter AI architecture.

Pega's Predictable AI model reframes the token question entirely. Rather than asking "How many tokens did this cost?" it asks, "What outcome did this achieve?" Pricing is anchored to the completed unit of work – the case – not the computational path taken to complete it.

This means:

  • No token-based billing surprises. Costs are fixed for the duration of the contract.
  • No usage caps. Teams can use as much AI as the work requires without watching a meter.
  • Full governance. Because runtime logic remains deterministic, workflows are auditable, compliant, and traceable – even when AI does the heavy lifting at design time.

    The Predictable AI model is built for enterprises that need AI to be both powerful and governable – not a trade-off between capability and cost control.

Frequently asked questions on AI tokens

Token counts determine AI costs. Every interaction is billed at tokens consumed × price per token. In agentic systems, costs compound fast: Agents reread context at each reasoning step, creating the AI token tax that can make workflows up to 10X more expensive than deterministic equivalents. Controlling token usage means controlling AI spend.

The model truncates the oldest content. In live conversations, it "forgets" earlier exchanges. In automated workflows, output quality silently degrades – often without flagging an error.

You can estimate AI token usage by measuring the tokens in your prompts and AI-generated responses. Since token counts vary by model and workload, using an AI token calculator provides a more accurate estimate. Use Pega's AI Token Calculator to estimate token usage and better understand potential AI costs.

Ready to learn more?

AI Orchestration Intro carousel

Tech knowledge

Discover why businesses are increasingly relying on AI orchestration for better decision-making and business outcomes.
AI Automation carousel

Tech knowledge

Learn how AI governance helps ensure that AI systems are developed and used responsibly and ethically across the enterprise.
Legacy modernization slide show

Tech knowledge

Discover how agentic AI can autonomously plan, learn, and adapt – driving agile workflows across diverse applications.
AI Workflow automation

Tech knowledge

Discover how AI workflow automation is revolutionizing business operations by boosting organizational efficiency.
Legacy Application Modernization carousel

Tech knowledge

Discover how business process orchestration reduces manual tasks and eliminates bottlenecks, leading to faster processing times.
Intelligent automation slideshow

Tech knowledge

Learn how responsible AI builds trust, mitigates risks, and ensures ethical, fair, and beneficial AI adoption.

Cut AI cost with Pega Blueprint®

Learn more
Blueprintを作成する準備はできましたか?
ニーズに合った変革エンジンを選択してください。
ワークフローおよびアプリ設計向け

プロセスを再設計し、あらゆるワークフローを構築可能なアプリケーションに自信を持って変換します。

Pega Blueprint®
PEGA BLUEPRINT™
マーケティングおよびCX戦略の設計

すべてのタッチポイントで顧客ジャーニーとエンゲージメント戦略を可視化し、活性化します。

Pega Customer Engagement Blueprint™
シェアする Xで共有 LinkedInで共有 Copying...