Skip to content
AlphaCorp AI
Wave of light particles flowing through faint circuit traces on a dark background
AI & Data Fundamentals20 min read

What is tokenization?

Ignas Vaitukaitis, Founder & CEO of AlphaCorp AI

AI Agent Engineer ·

What is tokenization?
On this page(14)
  1. What is tokenization? The first step in every language model pipeline
  2. Why subword tokenization beat words and characters
  3. The three main tokenization algorithms: BPE, WordPiece and SentencePiece
  4. How GPT, Llama, Gemini and Claude tokenize text today
  5. Why tokenization decides what you pay and how much context you get
  6. Tokenization explains LLM failures in counting, arithmetic and safety
  7. Tokenization beyond text: images as patches, and the move past fixed vocabularies
  8. Frequently asked questions about tokenization
  9. What is a token in AI?
  10. How many tokens is a word?
  11. Is tokenization the same as encryption or data tokenization in payments?
  12. How do I count tokens for ChatGPT or Claude?
  13. Why do LLMs struggle to count letters in a word?
  14. How to put tokenization knowledge to work

Tokenization is the step that converts raw text into a sequence of discrete units called tokens, each mapped to an integer ID, before a language model computes anything. Every prompt, every document, every bill you receive from an AI vendor passes through it first. This explainer covers how the main algorithms work, how GPT, Llama, Gemini and Claude split text differently, why that sets your API costs, and which well-known LLM failures trace straight back to it. As of October 02, 2026, tokenizer design is shifting faster than at any point since GPT-2.

  • Llama 3’s 2024 tokenizer compresses English at 3.94 characters per token, up from 3.17 under Llama 2 (Meta, “The Llama 3 Herd of Models”)
  • Claude Opus 4.7 and later models produce roughly 30% more tokens for the same text than earlier Claude models (Anthropic token-counting documentation, 2026)
  • Tokenized length for identical content differs by up to 15 times between languages (Petrov et al., NeurIPS 2023)
  • Inserting commas to align digit chunks lifted GPT-4’s large-number addition accuracy from 84.4% to 98.9% (Singh and Strouse, 2024)
  • Meta’s Byte Latent Transformer matched tokenizer-based model quality while cutting inference compute by up to 50% (Pagnoni et al., ACL 2025)

What is tokenization? The first step in every language model pipeline

Tokenization is the process of converting raw text into a sequence of discrete units called tokens, each mapped to an integer ID, before a language model does any computation at all. The model never sees your letters or words. It sees a list of numbers, and each number is looked up in an embedding table that turns it into a vector the neural network can work with. Everything the model “knows” about your prompt passes through this one step first.

That makes it the most consequential piece of code most people never look at.

OpenAI’s own documentation gives the cleanest demonstration of how strange the results can be. The string "ChatGPT is great!" splits into six tokens: Chat, G, PT, is, great, and !. A near-identical string, "tiktoken is great!", breaks into t, ik, token, is, great, ! under the GPT-4 encoding, according to OpenAI’s tiktoken cookbook on counting tokens.

Input stringTokens producedCount
ChatGPT is great!Chat / G / PT / is / great / !6
tiktoken is great!t / ik / token / is / great / !6

Two things in that table catch people off guard the first time they paste text into a tokenizer viewer. The space is part of the token. is and great carry their leading space with them, which means the same word at the start of a line and in the middle of a sentence can be two different tokens with two different IDs. And the splits ignore meaning: “GPT” shatters into G and PT, while “token” survives whole inside a word the tokenizer has clearly never treated as a unit.

So what counts as a token? Any of the following, depending on the scheme and the text:

  • A whole word (Chat, great)
  • A fragment of a word (ik, PT)
  • A single character (G, t)
  • A punctuation mark (!)
  • A raw byte, in byte-level schemes that can represent any string at all

Anyone who has sized chunks for a retrieval pipeline has run into the practical side of this: context limits are counted in tokens, and a word count only gets you a rough guess. A good rule of thumb for English is that a token is shorter than a word, often a good deal shorter for code, names, and anything with unusual spelling. The precise ratio depends entirely on which tokenizer the model uses.

Why subword tokenization beat words and characters

Subword tokenization won because it fixes the two failures that sank its predecessors: word-level tokenizers can’t represent any word they didn’t see during training, and character-level tokenizers produce sequences so long and so low in meaning that models learn from them poorly.

Start with words. A word-level tokenizer assigns one ID per word, which sounds natural until you ask what happens to a typo, a new product name, or a sentence in a language absent from the training set. There’s no ID. The word is out-of-vocabulary, and the model has no way to represent it. Worse, the vocabulary has no natural ceiling, because the set of possible words keeps growing with the data.

Characters go to the opposite extreme. Every string can be represented with a tiny vocabulary, so nothing is ever out of vocabulary. The price is that sequences become very long, and each unit carries almost no meaning on its own. The letter “t” tells the model nothing about whether “tokenization” or “tiger” is coming. Alqahtani and colleagues, in their EACL 2026 paper arguing that tokenizers are core design decisions in large language models, describe this as a worse inductive bias for the model.

AlphaCorp AIonline
Let's talk

Curious what AI could do for your business?

No jargon and no hard sell. Just a friendly look at where AI fits, and where it doesn't.

View Services

Subword schemes sit in the middle. Frequent words stay as single tokens. Rare or novel words get decomposed into pieces the model has seen before, all drawn from a vocabulary of fixed, finite size chosen in advance.

ApproachVocabulary sizeUnseen wordsSequence lengthMeaning per token
Word-levelUnbounded, grows with the dataFail (out-of-vocabulary)ShortHigh
Character-levelTinyAlways representableVery longNear zero
SubwordFixed, chosen in advanceSplit into known piecesModerateMixed

Subword schemes “solve the out-of-vocabulary problem while constraining the vocabulary size” compared to word-level models, and “improve the inductive bias of models while producing significantly shorter sequences” compared to character-level ones. (Alqahtani et al., EACL 2026)

One caveat the tidy table hides. The compromise is lopsided in practice, because which words count as “frequent” depends on the training corpus. A tokenizer built mostly on English keeps English words whole and fragments nearly everything else.

The three main tokenization algorithms: BPE, WordPiece and SentencePiece

The three algorithms behind nearly every production tokenizer are Byte Pair Encoding (BPE), which merges the most frequent symbol pairs; WordPiece, which merges the pairs that most increase training-data likelihood; and SentencePiece, which trains directly on raw sentences so it works for languages without spaces between words.

BPE is the oldest and the most widely used. The technique began life as a 1994 data-compression algorithm, and it was Sennrich, Haddow and Birch who in 2015 showed that segmenting text with the byte-pair-encoding compression algorithm let neural machine translation systems handle rare words (proper nouns, compounds, loanwords) as sequences of smaller, previously seen units. Their study reported gains of 1.1 BLEU on English-German and 1.3 BLEU on English-Russian translation. The training loop is almost embarrassingly simple:

  1. Split the training text into individual characters (or bytes).
  2. Count every pair of adjacent symbols.
  3. Merge the most frequent pair into a single new symbol and add it to the vocabulary.
  4. Repeat steps 2 and 3 until the vocabulary reaches its target size.
A four-step vertical process diagram of Byte Pair Encoding training. Step 1: split the training text into individual characters, or into raw bytes in byte-level variants. Step 2: count every pair of adjacent symbols across the corpus. Step 3: merge the most frequent pair into a single new symbol and add it to the vocabulary. Step 4: repeat the counting and merging until the vocabulary reaches its target size. The technique began as a 1994 compression algorithm and was adapted for neural machine translation in 2015, with reported gains of 1.1 BLEU on English-German and 1.3 BLEU on English-Russian.
Adapted from a 1994 compression algorithm, this loop gave neural machine translation gains of 1.1 BLEU on English-German and 1.3 BLEU on English-Russian. Source: Sennrich, Haddow and Birch, 2015.

OpenAI’s GPT-2 in 2019 popularized a byte-level variant for large language models. Starting from raw bytes means no string is ever unrepresentable, and the GPT-2 paper describes the goal as combining the empirical benefits of word-level models with the generality of byte-level approaches. One detail from that design still shapes every tokenizer viewer you’ll ever use: GPT-2 blocked merges across character categories (letters with punctuation, say), with an exception for spaces. That exception is why a leading space rides along with the word that follows it.

WordPiece came out of Google and powers BERT and its relatives. Mechanically it looks a lot like BPE. The difference is the merge criterion: where BPE picks the pair that appears most often, WordPiece picks the pair whose merge most raises the likelihood of the training data, as described in Google’s “Fast WordPiece Tokenization” paper at EMNLP 2021. At inference time it segments new text greedily, taking the longest vocabulary entry that matches at each position (maximum matching).

SentencePiece solved a problem the other two inherited. Both BPE and WordPiece in their original forms assumed the input had already been split on whitespace, which quietly breaks for Japanese, Thai, and any other language that doesn’t put spaces between words. Kudo and Richardson’s 2018 tool trains subword models directly from raw sentences, yielding what they call “a purely end-to-end and language independent system.” It supports BPE-style merging as well as a probabilistic Unigram Language Model algorithm.

AlgorithmOriginMerge criterionAssumes whitespace pre-split?
BPE1994 compression; adapted for NMT in 2015Most frequent adjacent pairYes, in its original form
WordPieceGoogle, for BERTLargest gain in training-data likelihoodYes
SentencePieceKudo and Richardson, 2018BPE or Unigram LM, trained on raw textNo

Here’s the part I find underappreciated. These three algorithms are decades of work on a problem that most engineering teams treat as a solved library call. The choice between them, and the corpus they’re trained on, decides how your model sees every character it will ever read.

How GPT, Llama, Gemini and Claude tokenize text today

As of October 2026, the four major model families tokenize text with four different designs: OpenAI’s GPT models use the open-source tiktoken BPE library with vocabularies of roughly 100K to 200K entries, Meta’s Llama 3 uses a 128,256-token tiktoken-style vocabulary, Google’s Gemini and Gemma share a 262,000-entry SentencePiece tokenizer, and Anthropic’s Claude uses an unpublished tokenizer that developers can only query through an API endpoint. Vocabulary design has turned into a headline feature of model releases.

OpenAI ships several named encodings inside tiktoken. The cookbook lists o200k_base for GPT-4o and GPT-4o-mini, cl100k_base for GPT-4-turbo, GPT-4, GPT-3.5-turbo and the embedding models, and the legacy p50k_base and r50k_base encodings for Codex and GPT-3-era models. The jump from the roughly 100K-entry cl100k_base to the roughly 200K-entry o200k_base, which arrived with GPT-4o in May 2024, doubled the vocabulary to compress non-English text and code more tightly.

Meta made the single biggest leap. Llama 2 used a 32K-entry SentencePiece vocabulary. Llama 3, released in 2024, switched to a 128,256-token vocabulary built from 100K tiktoken-style tokens plus 28K tokens added for non-English coverage. The Llama 3 paper puts English compression at 3.94 characters per token, up from 3.17 under Llama 2, and Meta’s release post claims the new tokenizer produces up to 15% fewer tokens than Llama 2’s on comparable text. Fewer tokens per document means the model reads more text for the same training compute.

Google went larger still. Gemini and Gemma share a SentencePiece tokenizer with 262,000 entries, split digits, preserved whitespace and byte-level fallback. The Gemma 3 technical report from March 2025 calls the design more balanced for non-English languages, with better compression for Chinese, Japanese and Korean at a small token-count penalty for English and code.

Anthropic is the outlier. The company has never published Claude’s vocabulary or merge rules, and its documentation tells developers to count tokens through the Messages API endpoint rather than a third-party tool like tiktoken. The same documentation, as of October 2026, states that Claude Opus 4.7 and later models (including Claude Fable 5.1 and Claude Mythos 5.1) use a newer tokenizer under which the same input produces roughly 30% more tokens than on earlier Claude models.

Model familyAlgorithm and libraryVocabulary sizeDistinctive choice
OpenAI GPT-4oBPE, tiktoken o200k_base~200K (2024)Doubled from cl100k_base for non-English and code
Meta Llama 3BPE, tiktoken-style128,256 (2024)28K tokens added for non-English coverage
Google Gemini / Gemma 3SentencePiece262,000 (2025)Split digits, byte-level fallback
Anthropic Claude Opus 4.7+UnpublishedUnpublished~30% more tokens per input than earlier Claude models (2026)

One pattern stands out. Three vendors expanded their vocabularies so that text compresses into fewer tokens, while Anthropic’s 2026 change moved the token count for identical text in the opposite direction. Both are legitimate engineering choices. Neither is something you can see from the outside without pasting your own text into the vendor’s counter.

Why tokenization decides what you pay and how much context you get

Tokenization decides what you pay and how much context you get because commercial APIs bill per token and cap context windows in tokens, which makes the tokenizer the exchange rate between the text you send and the money and context it consumes. Anthropic’s pricing and rate-limit documentation ties directly to counts from its count-tokens endpoint, and OpenAI frames its token-counting tools as cost-estimation instruments.

The rate moves. When a vendor changes tokenizers, every budget built on the old count is wrong on the day the new model ships. A tokenizer that yields roughly 30% more tokens for the same text, with per-token prices unchanged, means roughly 30% more cost per request and roughly 30% less of your document fitting in the window. Teams that pin spend to word counts discover this on the invoice.

The rate also differs by language, and this is where the economics stop being a bookkeeping matter. Vocabularies are trained on corpora that lean heavily toward English and a few other high-resource languages, so an English sentence stays in big tokens while the same sentence in Telugu gets shredded into many small ones. Researchers call this the token tax.

  • Petrov, La Malfa, Torr and Bibi (NeurIPS 2023) measured tokenization length differences between languages of up to 15 times for the same content, even for tokenizers explicitly trained to be multilingual, and argued this “induces unfair treatment for some language communities” in API cost, latency and effective context length.
  • Ahia and colleagues (EMNLP 2023) found that prompting and generation in languages such as Telugu and Amharic can cost up to 4 times more than English under commercial pricing.
  • A 2025 study titled “The Token Tax” documented the same disparity as systematic bias in multilingual tokenization.
  • A September 2026 paper proposed a formal “token-cost ledger” to separate the part of the tax that better engineering can remove from the part that is inherent to how a language is written.
  • MAGNET, a 2024 method, attacks the problem at training time with adaptive gradient-based tokenization aimed at multilingual fairness.
Two headline figures on the multilingual token tax. First, up to 15 times difference in tokenized length between languages for identical content, measured by Petrov and colleagues at NeurIPS 2023 and found even in tokenizers trained to be multilingual. Second, up to 4 times higher cost for prompting and generation in languages such as Telugu and Amharic compared with English under commercial pricing, measured by Ahia and colleagues at EMNLP 2023.
Identical content can take up to 15 times more tokens depending on the language, and prompting in Telugu or Amharic can cost up to 4 times more than in English. Source: Petrov et al., NeurIPS 2023; Ahia et al., EMNLP 2023.

Our team works across English, Portuguese and Spanish every day, and the gap shows up in ordinary work: run a Portuguese support transcript and its English translation through the same tokenizer viewer and the Portuguese version comes back longer, which means it costs more per call and leaves less room for retrieved context in the same window. The 2026 ledger paper’s distinction matters here. Some of that gap is a design choice vendors can fix. Some of it is the alphabet.

Three years of research, from 2023 through 2026, say the tax has narrowed in places but has not gone away.

Tokenization explains LLM failures in counting, arithmetic and safety

Tokenization explains a whole class of LLM failures because a token is usually an opaque chunk of several characters, so the model never receives the individual letters or digits inside it, and because the same text can be split more than one way, including ways that safety training never saw. The famous mistakes look like reasoning errors. Most of them are input errors.

Letter counting is the simplest case. If “strawberry” arrives as a single token, the model was never handed the letters r, r and r to count. Alqahtani and colleagues’ 2026 analysis of tokenization behavior describes exactly this mechanism, and it generalizes to spelling backwards, rhyming and any task that needs character-level access the vocabulary hid.

Arithmetic fails for a sharper version of the same reason. Singh and Strouse showed in their 2024 study “Tokenization counts” that GPT-3.5 and GPT-4-era tokenizers group digits into chunks of up to three reading from the left. Standard arithmetic carries from the right. So the number 1234567 might tokenize as 123 456 7 while the addition the model needs lines up as 1 234 567. Their fix was almost comically cheap: inserting commas to force aligned chunks raised GPT-4’s addition accuracy on large numbers from 84.4% to 98.9%. That is a prompt-formatting decision worth a 14-point accuracy gain, and most teams never test it.

A bar chart of GPT-4 accuracy on large-number addition under two prompt formats. With commas inserted to align digit chunks, accuracy reaches 98.9%, the highlighted bar. Without comma alignment, under default tokenization, accuracy is 84.4%. The gap is 14.5 percentage points, measured by Singh and Strouse in 2024.
Inserting commas to force aligned digit chunks raised GPT-4’s large-number addition accuracy by more than 14 points, from 84.4% to 98.9%. Source: Singh and Strouse, 2024.

Llama, PaLM and Mistral avoid the problem by design: every digit becomes its own token. Liu and Low’s 2023 Goat paper credits this digit-level tokenization as part of why a fine-tuned LLaMA model beat GPT-4 on large-number arithmetic. One caution on all of these numbers: the published measurements come from 2023 and 2024-era tokenizers, and whether later OpenAI encodings changed digit handling is something the vendor’s own counter will tell you faster than any paper.

Safety is the uncomfortable one. A 2025 paper on “Adversarial Tokenization” found that models retain semantic understanding of non-canonical tokenizations of a sentence even though their safety alignment was trained only on the canonical split. An attacker can re-segment a harmful request, changing zero characters, and slip past a filter that was looking at the characters. A related 2024 study, “Tokenization Matters!”, degraded model outputs on otherwise-normal input the same way.

FailureMechanismPublished evidence
Miscounting lettersMulti-character tokens hide individual lettersAlqahtani et al., 2026
Wrong addition on large numbersLeft-to-right digit chunks misalign with right-to-left carryingSingh and Strouse, 2024: 84.4% to 98.9% with commas
Safety filter bypassNon-canonical splits understood but never covered by alignment“Adversarial Tokenization”, 2025
Subtly wrong completionsPrompt’s final token boundary differs from what the model would generate“Sampling from Your Language Model One Byte at a Time”, 2025

The last row is the one that bites developers who think they’ve done everything right. The prompt boundary problem, described in a 2025 paper on byte-level sampling, happens when your prompt ends mid-token from the model’s point of view, so the first generated token has to start from a boundary the model would never have chosen itself. The standard fix is token healing: back up over the trailing token and let the model regenerate it. Without it, a prompt ending in http: or a half-typed identifier produces completions that look fine and are quietly off.

Tokenization beyond text: images as patches, and the move past fixed vocabularies

Tokenization applies to more than text: Vision Transformers cut an image into fixed-size patches and project each one into a token, and the newest research direction drops the fixed vocabulary altogether in favor of byte patches sized by how predictable the next byte is.

Images first. A Vision Transformer divides a picture into a grid of equal-size patches, each linearly projected into a token that stands for one local region, as Google Research describes in its post on learning to tokenize images for Vision Transformers. The same post makes a point that carries straight back to language: the quality and allocation of those visual tokens directly set model quality. Same idea, different input. Continuous raw data becomes discrete, learnable units, and the way you draw the boundaries decides what the model can see.

Now the text frontier. Meta’s Byte Latent Transformer (BLT), published at ACL 2025 by Pagnoni and colleagues, removes the fixed tokenizer entirely. It reads raw bytes and groups them into patches whose size changes with the entropy of the next byte. Predictable stretches, like the tail of a common word, collapse into long patches that get little compute. Surprising stretches get short patches and more.

The results are why people are paying attention:

  • Scale tested: a FLOP-controlled scaling study up to 8 billion parameters and 4 trillion training bytes (2025)
  • Quality: matched tokenizer-based models at fixed cost
  • Inference compute: cut by up to 50%
  • Noisy and adversarial input: handled better than standard tokenized models
Three headline figures from the Byte Latent Transformer study published at ACL 2025. Inference compute cut by up to 50% compared with tokenizer-based models, with quality matched at fixed cost, the highlighted figure. Models tested up to 8 billion parameters in a FLOP-controlled scaling study. Training used up to 4 trillion raw bytes, with no fixed vocabulary.
Meta’s Byte Latent Transformer matched tokenizer-based model quality at fixed cost while cutting inference compute by up to 50%, tested up to 8 billion parameters and 4 trillion training bytes. Source: Pagnoni et al., ACL 2025.

The robustness result follows from the design. With no vocabulary, there is no canonical segmentation for an attacker to exploit.

Whether BLT or a successor displaces subword vocabularies in production is open. The 2025 study stops at 8B parameters, and frontier models run far larger. What has already changed is the field’s attitude. Alqahtani, Nayeem, Laskar, Mohiuddin and Bari argue in their EACL 2026 paper that tokenization deserves the same scrutiny as architecture and training objective, and a September 2026 survey describes it as a decision that “shapes sequence length, computational cost, multilingual equity, evaluation, and security.”

Tokenization has long been “treated as an afterthought, implemented with insufficient scrutiny” and should instead be “an essential component of model development.” (Alqahtani et al., EACL 2026)

I’d put it more bluntly. If you haven’t inspected how a model splits your own data, you don’t yet know what that model reads.

Frequently asked questions about tokenization

These are the questions people most often search about tokenization, each answered in a few sentences that stand on their own.

What is a token in AI?

A token is the smallest unit of text a language model processes: a whole word, part of a word, a single character, a punctuation mark or, in byte-level schemes, a raw byte. Each token maps to an integer ID that the model looks up in an embedding table before any computation. OpenAI’s example string “ChatGPT is great!” becomes six tokens, with “GPT” split into “G” and “PT”.

Built for production

What could a custom AI agent take off your plate?

We build production-grade AI systems that quietly handle the busywork, so your team can focus on the work that actually matters.

View Services

How many tokens is a word?

There is no fixed answer, because the ratio depends on the tokenizer and the language. For English under Llama 3’s tokenizer, Meta’s 2024 paper measured 3.94 characters per token, so an ordinary English word runs to roughly one token while rare words, code and names split into several. Text in lower-resource languages can take many times more tokens for the same meaning.

Is tokenization the same as encryption or data tokenization in payments?

No. In payments and data security, “tokenization” means replacing sensitive data such as a card number with a stand-in value so the original stays protected. In language models, tokenization is plain text segmentation: splitting a string into units a model can process, with nothing concealed. The two share a name and little else.

How do I count tokens for ChatGPT or Claude?

Use the vendor’s own counter. For OpenAI models, the open-source tiktoken library with the matching encoding (o200k_base for GPT-4o) returns exact counts offline. For Claude, Anthropic publishes no tokenizer and directs developers to the Messages API count-tokens endpoint, which as of October 2026 is the only accurate method. Third-party calculators are estimates for Claude and go stale whenever a vendor changes tokenizers.

Why do LLMs struggle to count letters in a word?

Because the model never sees the letters. A common word like “strawberry” usually arrives as a single token, an opaque ID with no internal structure, so the model has to recall the spelling from training instead of inspecting it. Alqahtani and colleagues’ 2026 analysis traces letter-counting failures to this mechanism. Long-number arithmetic breaks for a related reason when digits are grouped into multi-digit tokens.

How to put tokenization knowledge to work

Putting tokenization knowledge to work comes down to five habits that cost almost nothing and save real money and real accuracy.

  1. Count tokens with the vendor’s own tool before you estimate cost or size a chunk. Word counts are guesses.
  2. Re-benchmark every token budget when a model you depend on changes tokenizer. Published vendor changes have moved counts for identical text by 15% to 30%.
  3. Run your prompts and documents through the counter in every language your users write, and budget for the longest.
  4. Format large numbers with separators when a model must do arithmetic on them, then verify the gain on your own cases.
  5. Treat tokenizer choice as a design decision when you pick or fine-tune a model, with the same scrutiny you give architecture.

The tokenizer is already shaping your results and your bill. Whether you’ve looked is the only variable. If you want a second pair of eyes on how your current models split, bill and chunk your data, AlphaCorp AI’s AI integration audit starts with exactly that count.

Share
Newsletter · Weekly

Stay Ahead of AI

One email per week with the AI engineering insights, agent builds, and tools that actually matter.

No spamUnsubscribe anytimeFree forever

In every issue
  1. 01One agent build, taken apart step by step
  2. 02The tools that earned a place in our stack this week
  3. 03What broke in production, and what we changed

Written by Ignas Vaitukaitis, founder of AlphaCorp AI.

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor

Ready to Ship
Your AI System?

Book a free call and let's talk about what AI can do for your business. No sales pitch, just a real conversation.