GitHub Repo

Token lab / 01

Tokenization

Deep dive for AI engineers

Think of tokenization as a translator. You write normal text, such as Hello world, and the tokenizer turns it into numbers like [101, 7592, 2088, 102]. The model only ever sees those numbers.

Those numbers come from a fixed vocabulary — simply a list of all the pieces the tokenizer knows. Here, 101 and 102 mark the start and end of the sentence, while 7592 and 2088 stand for the pieces in between. When the model finishes, the tokenizer runs in reverse and turns the numbers back into readable text.

Tokenizers such as BPE, WordPiece, and SentencePiece are the recipes for choosing those pieces — and for reversing the process afterwards.

Never seen this before? Start here.

A language model cannot read. It can only do arithmetic. So every piece of text has to be turned into numbers before the model sees it, and turned back into text afterwards. Tokenization is that conversion.

The only real question in this whole topic is: how big is one piece? A whole word? Half a word? One letter? One byte? Every answer trades something away, and this page lets you feel each trade in a live playground instead of taking it on trust.

Suggested path if you are new: 02 how it works → 03 Unicode & UTF-8 → 04 play with the splitter → 07 how BPE works → 08 train one yourself. Sections 09–11 explain why real models chose bytes.

How tokenization works

From Text to Numbers

A neural network works with numbers, not with words or sentences. So before a transformer can process text and predict what comes next, the text must first be converted into a sequence of token IDs.

For example, a BERT-style tokenizer might turn Hello world into [101, 7592, 2088, 102]. Here, 101 and 102 are special tokens that mark the beginning and end of the sequence, while 7592 and 2088 represent the pieces of text in between.

The transformer does not receive the original words. It receives these IDs, which are then turned into numerical vectors called embeddings. After the model produces its output, the tokenizer is used again to turn token IDs back into readable text.

At a high level — the full round trip
  1. Text Hello world
  2. Tokenizer encode
  3. Tokens hello · world
  4. Token IDs 7592 · 2088
  5. Embeddings vectors
  6. Transformer the model
  7. Output IDs predicted
  8. Tokenizer decode
  9. Text readable again

what a human reads the tokenizer, used twice numbers only the model

The one takeaway. The tokenizer is the bridge between human-readable text and the numbers a neural network can process. It is the only part of the stack that speaks both languages, which is why it appears at both ends of the diagram.

Example: if two tokenizers both see Hello world, they may agree on the idea of “hello” and “world” but still assign different IDs. The number is just a label inside that tokenizer’s own vocabulary.

The Tokenization Pipeline

Tokenization is not just "splitting a sentence into words." A tokenizer usually performs several steps, and each step has a specific job.

Think of the process as a pipeline. Each stage prepares the text for the next one. If a stage is skipped or done badly, you can get wrong lengths, unknown tokens, or broken multilingual text.

Example: "Hello, I'm learning AI!" may first stay as raw text, then be normalized, then split into pieces like Hello, ,, I, ', m, learning, AI, ! before those pieces are mapped to IDs.

Here is the full path from text to model input and back again:

01 Raw text Unicode string
02 Normalize lowercasing, Unicode NFKC
03 Pre-tokenize split on spaces / punctuation
04 Model word, subword, char, or byte
05 IDs integers the net sees
06 Decode IDs back to text

The granularity of step 4 is the main design choice. Each level answers one question: what is one token?

Example: the word lowest can be a single word token, two subword pieces like low + est, six characters, or six bytes before any merges happen. The answer depends on the tokenizer you choose.

Word

One token ≈ one whitespace word. Tiny sequences, huge vocabulary, and any new word becomes unknown. Fine for small closed domains. Fragile for open web text.

Example: I love football becomes three tokens, but ChatGPT may become [UNK] if the word was never trained.

Subword

One token is a frequent chunk: token + ##ization, or low + er. This is how BERT, GPT, T5, and Llama actually work. Three training recipes live in the algorithm labs below.

Example: playing might become play + ing, so the model can reuse pieces it already knows instead of learning every word from scratch.

Character

One token is a Unicode code point (sometimes a grapheme cluster). Vocabulary stays small. Sequences get long. Spelling and morphology become the model’s problem.

Example: hello becomes h, e, l, l, o. A short word looks harmless, but a whole sentence gets much longer very quickly.

Byte

One token is a UTF-8 byte (0–255). Nothing is unknown — every Unicode string is some byte sequence. GPT-2-style BPE starts here, then merges bytes into longer pieces.

Example: A is 1 byte, é is 2 bytes, and 👋 is 4 bytes before merges. The model always has a way to spell the text.

Rule of thumb. Coarser tokens → shorter sequences, bigger vocab, more out-of-vocabulary pain. Finer tokens → longer sequences, smaller vocab, more compute per sentence. Subword is the industry compromise.

Example: a word tokenizer may use one token for house, a subword tokenizer may use house or hous + e, a character tokenizer uses five tokens, and a byte tokenizer starts with the UTF-8 bytes underneath it.

What a vocabulary actually is

A vocabulary is a two-way map: piece ↔ integer. Special pieces almost always sit at the front: padding, unknown, start, end, mask. Training a tokenizer means choosing the piece list so that a large corpus compresses well. Using a tokenizer means looking pieces up — you do not invent new IDs at inference time.

Token
One atomic piece after splitting. Might be a word, a subword, a character, or a byte.
Type / ID
The integer that names that piece inside the vocabulary. The model embeds this ID.
OOV / UNK
Out of vocabulary: a piece the vocab never listed. Word-level tokenizers hit this constantly.
Detokenize
Glue pieces back into a string. Spaces, ##, and the SentencePiece ▁ all exist to make this reversible.

Leave when: you can name the six pipeline stages and say which one chooses word vs subword vs byte.

Unicode and UTF-8

Before a tokenizer ever runs, your text is already numbers. Two different standards did that job, and beginners almost always blur them together. Separating them makes everything after this page easier.

Unicode: the catalogue of characters

Unicode is a universal standard that assigns a unique code point — a number — to every character. It is a giant numbered catalogue that everyone in the world agrees on:

  • A → U+0041 (decimal 65)
  • é → U+00E9 (decimal 233)
  • আ → U+0986 (decimal 2438)
  • 😊 → U+1F60A (decimal 128522)

The U+ prefix just means “this is a Unicode code point, written in hex.” Because every language, symbol, and emoji has its own agreed number, a document written in Tokyo opens correctly in Zürich. That is the whole point of Unicode: consistent identity for characters.

Example: the letter A is always U+0041, whether it appears in a sentence, a filename, or a code comment. Unicode gives the character its identity before any tokenizer decides what to do with it.

Crucially, Unicode says nothing about how those numbers are stored in a file or sent over a network. It only hands out the numbers.

UTF-8: the rule for storing those numbers as bytes

UTF-8 is an encoding: it converts Unicode code points into bytes so text can be stored on disk or transmitted. It uses a variable number of bytes per character:

  • ASCII characters (U+0000–U+007F) use 1 byte
  • Everything else uses 2, 3, or 4 bytes
The relationship in one line. Unicode defines which characters exist and what number each one has → UTF-8 defines how that number is written down as bytes. Unicode is the catalogue; UTF-8 is the packing rule.

Example: A is one byte in UTF-8, but 😊 is four bytes. That is why byte-level tokenizers see more pieces for emoji and many non-Latin scripts until later merges compress them again.

The same four characters, all the way down

Read this table left to right and you have followed one character from human symbol, to Unicode number, to the actual bytes on disk:

Character Unicode name Code point UTF-8 bytes (hex) Bytes
A LATIN CAPITAL LETTER A U+0041 41 1
é LATIN SMALL LETTER E WITH ACUTE U+00E9 C3 A9 2
আ BENGALI LETTER A U+0986 E0 A6 86 3
日 CJK UNIFIED IDEOGRAPH-65E5 U+65E5 E6 97 A5 3
😊 SMILING FACE WITH SMILING EYES U+1F60A F0 9F 98 8A 4

Notice that A costs one byte and 😊 costs four. That single fact explains the byte column in the four-way lab below, and it is why the same sentence can be cheap in English and expensive in Bengali or Japanese.

Example: café looks short on screen, but the accent means the code-point and byte counts are different from plain cafe. That is why the inspector below separates graphemes, code points, and bytes.

Build your own rows in the UTF-8 inspector — type anything and watch graphemes, code points, and bytes diverge.

Three different ways to count “length”

Beginners lose hours to this. The string café 👋 has three defensible lengths, and each one is correct for a different question:

Example: a browser may show 👋 as one visible character, but a string API can count code points differently, and the byte-level tokenizer always counts the UTF-8 bytes underneath it.

Grapheme
What a human calls “one character” on screen. é is one grapheme even when it is built from e + a combining accent.
Code point
One entry in the Unicode catalogue. é can be one code point (U+00E9) or two (U+0065 + U+0301) — which is exactly what normalization (NFKC in the pipeline) is for.
Byte
One UTF-8 byte, 0–255. This is what is actually stored, and what byte-level tokenizers consume.

How big is Unicode, really?

1,114,112 code points in the Unicode codespace (U+0000–U+10FFFF)
159,801 characters actually assigned as of Unicode 17.0 (Sept 2025)
172 scripts (writing systems) supported — serving hundreds of languages
256 possible values of one UTF-8 byte — and that never grows

Note the wording: 172 scripts, not 172 languages. One script such as Latin or Devanagari is shared by many languages. Also note that the assigned-character count grows with every Unicode release — the 256 byte values never do. Hold on to that contrast; section 10 turns it into a design decision.

Why UTF-8 won. It is backward compatible with ASCII (English text is byte-for-byte identical to a 1960s ASCII file), it has no byte-order ambiguity, it can encode every Unicode character, and it is self-synchronizing — you can jump into the middle of a stream and find the next character boundary. Python and JavaScript hold strings differently in memory (Python uses code points, JavaScript uses UTF-16 units), but files, HTTP, and model tokenizers overwhelmingly speak UTF-8.

Leave when: you can say in one sentence what Unicode gives you, what UTF-8 gives you, and why 😊 is four bytes but A is one.

Four-way lab

Type anything. The same string is split four ways at once. Watch how token count explodes as units get smaller, and how punctuation, emoji, and non-English scripts behave.

Example: try café 👋 or বাংলা 日本語. The word view stays short, the character view grows, and the byte view grows fastest before merges have a chance to compress repeated patterns.

Live splitter

Word

0

Naive subword

0

Character

0

UTF-8 byte

0

“Naive subword” here is a teaching trick: 3-character chunks with WordPiece-style ##. Real BPE / WordPiece / SentencePiece learn those chunks from data — open Algorithm labs.

Quick check: is the “Naive subword” pane real BPE?

Leave when: you can explain why the byte column is longer than the character column for 👋.

Word level

The oldest idea: split on whitespace (and usually punctuation). Each distinct word string gets its own ID. “cat” and “cats” are unrelated. “Tokenization” is a brand-new type even if the model already knows “token”.

Example: I love football becomes three IDs in a word tokenizer, but I love ChatGPT can fall apart if ChatGPT was never seen in training. That is the OOV problem in one sentence.

That is the out-of-vocabulary (OOV) problem. At training time you only store the words that appeared often enough. At test time a name, a typo, or a new product word becomes [UNK]. The model then sees a hole, not the spelling.

The OOV trap

Lookup

Word IDs after lookup

Try replacing dog with cats, a name, or a Bengali word. Word-level tokenization has no way to reuse letters it already knows. That single failure is why modern LLMs are not word tokenizers.

Leave when: you can say why cat and cats are unrelated IDs under word-level tokenization.

Subword level

Subword tokenization keeps frequent words intact and splits rare words into reusable pieces. unhappiness might become un + happiness, or un + happy + ness, depending on the algorithm and the corpus.

Example: playing can stay whole if it is common, but a rarer word like unhappiness is more useful when split into pieces that other words can reuse. That is why subword tokenization feels like a compromise: it keeps common words short and rare words possible.

Same corpus, three scoring rules. Play all three in the lab, then train the real libraries in Jupyter.

BPE
Merge the most frequent adjacent pair, again and again. GPT family, many Llama tokenizers.
WordPiece
Merge the pair that most improves a likelihood score, not raw count. BERT, DistilBERT. Continuation mark ##.
SentencePiece
Treat the raw string as the unit (spaces become the character ▁). Unigram LM or BPE. T5, ALBERT, many multilingual models.
Misconception to kill. BPE pieces are not guaranteed morphemes. est appears because the pair was frequent, not because English has a comparative suffix. Probe star vs widest with and without </w> in mind.

Example: a tokenizer might learn low + est from the corpus, then reuse those same pieces in lowest and widest even though the words are not built the same way in grammar textbooks. The pieces are learned from frequency, not from a dictionary of linguistics rules.

Leave when: you can name the scoring rule for each of the three algorithms in one sentence.

Which models use which tokenizer?

Before learning BPE, it helps to see where the main tokenizer methods are used. Vocabulary size means the number of token IDs the tokenizer can produce. It does not mean the number of words or ideas the model understands.

Model version Tokenizer method Starts with Vocabulary size
GPT-2 Byte-level BPE UTF-8 bytes 50,257
RoBERTa Byte-level BPE UTF-8 bytes 50,265
BERT base, uncased WordPiece Characters 30,522
ALBERT base v2 SentencePiece Unigram Characters 30,000
Llama 2 SentencePiece BPE Characters and bytes 32,000
Llama 3 Tiktoken-based BPE UTF-8 bytes 128,256
DeepSeek-V3 Byte-level BPE UTF-8 bytes 129,280
Qwen2.5 Byte-level BPE UTF-8 bytes 152,064
Kimi K2 Byte-level BPE UTF-8 bytes 163,840
Llama 3.3 70B on Groq Tiktoken-based BPE UTF-8 bytes 128,256
Always check the exact model version. Models in the same family can use different tokenizers. Vocabulary counts can also include special tokens such as [CLS], [SEP], or chat-control tokens, so sources sometimes report slightly different totals.

Why is Groq listed differently? Groq runs models from companies such as Meta, Google, and OpenAI on its inference platform. It does not give every hosted model one “Groq tokenizer.” A Llama model served by Groq still uses its Llama tokenizer, while a different hosted model uses its own tokenizer.

How to train your own tokenizer

You can train a tokenizer without training a language model. The tokenizer learns useful pieces from a collection of text called a training corpus.

  1. Collect representative text. Include the languages, code, names, numbers, and writing styles that your model will receive.
  2. Choose a method. Use BPE, WordPiece, or SentencePiece Unigram, depending on the behaviour you want.
  3. Choose a vocabulary size and special tokens. For example, request 32,000 token IDs and reserve tokens such as [UNK], [BOS], and [EOS] if your model needs them.
  4. Run the tokenizer trainer. It reads the corpus, counts patterns, and builds the vocabulary and rules. No neural-network training is involved yet.
  5. Test, save, and freeze it. Check common text, rare words, every target language, emoji, and code. Once model training begins, keep the token IDs fixed so their learned embeddings do not change meaning.
Try it in this project. The three Jupyter notebooks train BPE, WordPiece, and SentencePiece tokenizers on a small corpus, so you can inspect the learned vocabulary before using a large dataset.

How BPE works, step by step

Byte Pair Encoding is the algorithm behind the GPT family and most Llama tokenizers. It is also genuinely simple — simple enough to do by hand, which is what this section does before you drive it in the lab.

The algorithm

BPE is a greedy algorithm that builds a vocabulary by repeatedly merging the most frequent pair of adjacent tokens. “Greedy” means it takes the best pair available right now, with no lookahead and no going back to reconsider an earlier merge.

  1. Prepare the corpus. Split the text into words and count how often each distinct word occurs. Mark word boundaries — this lab uses an end-of-word symbol </w> — so merges can never run across two words.
  2. Start with individual characters as tokens. Every word is spelled out as a sequence of single characters. Those characters are the starting vocabulary. (Byte-level BPE starts from the 256 bytes instead — same algorithm, different atoms.)
  3. Count adjacent pairs. Across the whole corpus, count how often each pair of neighbouring tokens occurs, weighted by how often its word occurs.
  4. Merge the most frequent pair into one new token, everywhere it appears.
  5. Record the merge and add the new token to the vocabulary. The merge list is ordered — that order matters later.
  6. Repeat steps 3–5 until the vocabulary reaches its target size (or, equivalently, until you have performed a chosen number of merges). This is the stop condition.
Why BPE must stop merging

First, a merge means joining two neighbouring tokens and treating them as one new token. Suppose BPE starts with the word low split into three character tokens:

start          l + o + w       3 tokens
merge l + o    lo + w          2 tokens
merge lo + w   low             1 token

This is what “the word shrinks” means: the text does not lose letters, but the number of tokens becomes smaller. If BPE continued merging every possible pair, each word in the training corpus would eventually become one large token. For example, lowest would no longer be reusable pieces such as low + est; it would become the single token lowest. The vocabulary would then behave too much like a word-level vocabulary, with a separate large token for each training word.

Real BPE therefore stops at a limit chosen before training. That limit may be a target vocabulary size or a fixed number of merges. Think of it as a spending budget: BPE is allowed to create only a certain number of new tokens. GPT-2, for example, was given 50,000 merges. This keeps useful common pieces while avoiding a separate token for every complete word.

A second stopping rule can ignore very rare pairs. If a pair appears only once, creating a permanent vocabulary entry for it usually saves very little. This lab merges pairs with frequency 2 or more and stops when the best remaining pair has frequency below 2.

What happens after BPE stops? Training is finished. BPE saves the new vocabulary pieces and the merge rules in the exact order they were learned. These are then frozen: when the tokenizer receives new text, it does not count pairs or learn new merges. It simply starts from the smallest units and replays the saved rules that apply.

For example, imagine training stops after learning these four rules:

01  l  + o   → lo
02  lo + w   → low
03  e  + s   → es
04  es + t   → est

The saved vocabulary now includes pieces such as lo, low, es, and est. When BPE later sees lowest, it replays those rules and produces low + est. If it sees lower, it can produce low + e + r because no saved rule joins the final e and r. Each resulting piece is then looked up in the vocabulary and replaced with its token ID for the model.

Training chooses the pieces; encoding uses them. After the stopping point, the tokenizer no longer learns. It applies the frozen merge list to every new input and falls back to smaller known units whenever no larger saved merge applies.

For example, with the four saved rules above:

lowest  → low + est
lower   → low + e + r
lovely  → lo + v + e + l + y

lowest can use two large learned pieces. lower can use low, but its remaining letters stay separate because no e + r rule was learned. lovely can use only l + o → lo, so everything else falls back to individual characters. The tokenizer does not create a new lovely token while encoding it.

What is the disadvantage? Frozen rules make token IDs stable, but they also preserve the biases of the training corpus. Text similar to the training data gets large, efficient pieces; unfamiliar words, languages, or technical terms may be split into many small pieces.

  • Longer sequences: lowest uses 2 token positions, while lovely uses 5 in this example.
  • More computation: the transformer must process every token, so more pieces require more work and consume more of the context window.
  • Uneven efficiency: a language or subject that was rare in the training corpus may need far more tokens than common English text.
  • Difficult to update: adding new vocabulary pieces later changes the token-ID system and normally requires updating or retraining the model's embedding and output layers.

The text is still representable, so fallback is better than producing [UNK]. The disadvantage is mainly inefficiency, not loss of the original spelling.

BPE must know where one word ends and the next word begins. In the cat, it may join letters inside the, such as t + h → th, or inside cat, such as c + a → ca. But it must never join e + c, because e belongs to the and c belongs to cat.

The marker </w> means “this word ends here.” The sentence can be represented as t h e </w> and c a t </w>, so BPE knows not to merge across the boundary.

The starting pieces depend on the BPE version. Character-level BPE starts with letters such as c, a, and t. Byte-level BPE starts with byte values from 0 to 255. After that, both versions follow the same basic process: count neighbouring pairs, merge common pairs, and stop at the chosen limit.

Worked example: the classic four-word corpus

This is the corpus loaded in the lab below — four word types with their counts:

The × number tells you how many times that complete word appears in the training corpus. For example, low ×5 means the corpus contains five copies of low, while lower ×2 means it contains two copies of lower. The word is shown only once here to keep the table compact. spelled shows the starting character tokens, and </w> marks the end of the word.

low     ×5      spelled  l o w </w>
lower   ×2      spelled  l o w e r </w>
newest  ×3      spelled  n e w e s t </w>
widest  ×2      spelled  w i d e s t </w>

Step 3 counts every adjacent pair. The pair l+o appears in low (5 times) and in lower (2 times) — 7 occurrences, the highest in the corpus. So it wins the first merge. Repeat, and BPE produces this exact ordered merge list:

In the merge list, freq means the total number of times the adjacent pair appears across the corpus at that step. Word counts are included in that total. For example, l + o appears once inside each copy of low and lower, so its frequency is 5 + 2 = 7. After it is merged into lo, the pair lo + w also appears seven times. The frequency is recalculated after every merge because the available token pairs have changed.

01  l + o        →  lo          freq 7
02  lo + w       →  low         freq 7
03  low + </w>    →  low</w>     freq 5
04  e + s        →  es          freq 5
05  es + t       →  est         freq 5
06  est + </w>    →  est</w>     freq 5
07  n + e        →  ne          freq 3
08  ne + w       →  new         freq 3
09  new + est</w> →  newest</w>  freq 3
10  low + e      →  lowe        freq 2
11  lowe + r     →  lower       freq 2
12  lower + </w>  →  lower</w>   freq 2
13  w + i        →  wi          freq 2
14  wi + d       →  wid         freq 2
15  wid + est</w> →  widest</w>  freq 2

After merge 15 no pair is left with frequency ≥ 2, so training stops on the guard rather than the budget. Reproduce every line of this by dragging the slider in the BPE lab — the merge log there is generated by the same algorithm.

Encoding a word BPE has never seen

This is the payoff. The word lowest is not in the training corpus. To encode it, BPE spells it out in characters and then replays the learned merges in the order they were learned:

start        l  o  w  e  s  t  </w>
after 01     lo w  e  s  t  </w>        (l+o)
after 02     low   e  s  t  </w>        (lo+w)
after 04     low   es t  </w>           (e+s)
after 05     low   est   </w>           (es+t)
after 06     low   est</w>              (est+</w>)

result       ["low", "est</w>"]  →  2 tokens, decodes back to "lowest"

Merge 03 (low+</w>) does not apply, because in lowest the piece low is not followed by the end of the word. That is precisely why the </w> mark exists: it lets BPE distinguish “low as a complete word” from “low as the start of a longer word.”

Advantage and tradeoff. The advantage is total coverage: BPE can represent a word it has never seen by using smaller known pieces, so byte-level BPE never needs [UNK]. The disadvantage is inefficiency: an uncommon word may need many tokens, which uses more context space and computation. In the worst case, the tokenizer emits one character or one byte at a time.

The strawberry and blueberry problem

People see the shared ending berry immediately. BPE does not begin with that meaning. It only learns whichever neighbouring pieces were frequent in its training corpus. One tokenizer might produce:

strawberry  →  straw + berry       2 tokens
blueberry   →  blue  + berry       2 tokens

Another tokenizer, trained on different text or with a smaller vocabulary, might split the same words less neatly:

strawberry  →  str + aw + ber + ry     4 tokens
blueberry   →  blue + b + err + y      4 tokens

Both tokenizations are valid because every piece joins back into the original word. The second version is less efficient and does not expose berry as one reusable piece. It can also make questions such as “How many rs are in strawberry?” harder: the model receives token IDs for chunks, not a tidy list of individual letters. Tokenization contributes to that difficulty, although it is not the only reason a language model may answer a spelling question incorrectly. The exact split depends on the tokenizer, so always inspect the real tokens instead of assuming these example splits.

What BPE is not

  • Not a morphological analyser. est emerged because that pair was frequent, not because English has a superlative suffix. On a different corpus you get different pieces, and plenty of them straddle real morpheme boundaries.
  • Not optimal. Greedy means locally best at each step. There is no guarantee that this merge sequence is the best possible vocabulary of that size — only that it is cheap to compute and works well in practice.
  • Not the same at train and encode time. Training discovers the ordered merge list from counts. Encoding just replays that fixed list. No new tokens are ever invented at inference time.

Leave when: you can state BPE’s stop condition correctly, and explain why running the merge loop “until no more pairs appear” would defeat the purpose.

Algorithm labs

This lab lets you compare BPE, WordPiece, and SentencePiece on the same corpus (the text used to learn token pieces). Because all three trainers receive the same text, you can see how their different learning rules create different vocabularies and word splits.

Choose an algorithm, move the slider to control how much it learns, and enter a probe word to see how the trained tokenizer splits new text. The browser versions are simplified so you can inspect every step. The Jupyter notebooks later in the lesson repeat the experiments with the production libraries HuggingFace tokenizers and sentencepiece 0.2.2.

Subword trainers

Sennrich toy = four word types (low / lower / newest / widest) so WordPiece can form low + ##est. Course corpus matches data/tiny_corpus.txt used by HuggingFace / SentencePiece cells in the notebooks — merges will differ. That is expected.

Slider 0–20 = number of merge steps

Encoded pieces

0


          

Same probe, three encodings

Frozen to the corpus, slider, and probe above. Switch SentencePiece Unigram/BPE with the mode buttons — the third column follows.

Default corpus is the Sennrich toy set so WordPiece can form low + ##est. Paste the course corpus and watch WordPiece prefer rare pairs like qu — that is the likelihood score doing its job, and why production models train on huge text.

Ready to move on when: you can explain why BPE, WordPiece, and SentencePiece split the same word into different pieces. The tokenizers agree on the word; they just disagree on where to split it.

Character level

A character-level tokenizer usually gives each Unicode code point its own token ID. Its vocabulary contains the distinct characters found in the training text. An English-focused vocabulary may contain around one hundred letters, digits, punctuation marks, and special symbols. A multilingual vocabulary may need thousands.

However, one visible symbol is not always one code point:

  • JavaScript can count one emoji as two. JavaScript's "👋".length is 2 because length counts UTF-16 code units. Use Array.from(text) to iterate over code points. Python's for ch in text already iterates over code points.
  • One symbol can contain several code points. What a reader sees as one symbol is called a grapheme cluster. An accented letter can be stored as one code point, such as é, or as e followed by a separate accent code point. Therefore, café can contain four or five code points even though it looks identical on screen.

Try café, emoji, Bengali text, and a combining accent in the Four-way lab. Compare the character count with the UTF-8 byte count to see the difference yourself.

Ready to move on when: you can explain why 👋 looks like one symbol even though JavaScript reports "👋".length as 2.

Byte level

Computers store UTF-8 text as bytes. A byte is simply a number from 0 to 255, so there are exactly 256 possible byte values. A byte-level tokenizer begins with one token for each of those values. This small base vocabulary can represent any valid UTF-8 text, including unfamiliar names, emoji, code, and mixed languages. If no larger token matches, the tokenizer can always fall back to the original bytes instead of producing [UNK].

The disadvantage is that raw bytes can create long sequences. The letter A uses one UTF-8 byte, but 👋 uses four. Without any learned shortcuts, that one emoji would therefore require four byte tokens.

Traditional BPE vs byte-level BPE

Byte-level BPE does not replace the original BPE algorithm. Both versions begin with basic symbols, count adjacent pairs, merge the most frequent pair, and repeat. The difference is the starting unit: traditional character-level BPE learns from the Unicode characters found in its training data, while byte-level BPE first converts text to UTF-8 and starts from all 256 possible byte values.

That fixed byte vocabulary can represent every Unicode string, even when a particular emoji, Hindi character, or Chinese character never appeared during training. Instead of replacing unfamiliar text with [UNK] and losing information, the tokenizer falls back to bytes and then uses the same BPE merges to combine frequent byte sequences into larger tokens. This universal base vocabulary and efficient merge process are why GPT-style tokenizers use byte-level BPE.

Open the Byte-level BPE vs traditional BPE playbook in Colab.

What does “BPE on top of bytes” mean?

BPE starts with individual bytes and repeatedly joins byte sequences that appear together often. Think of each learned merge as a shortcut. For example, the word cat has these UTF-8 bytes, written in hexadecimal:

text               cat
UTF-8 bytes        63  61  74
starting tokens    [63] [61] [74]

learn 63 + 61      [6361] [74]       2 tokens
learn 6361 + 74    [636174]           1 token

The final token still represents the same three bytes; BPE has only grouped them into one frequently used piece. Decoding expands the piece back to 63 61 74, and UTF-8 turns those bytes back into cat.

The same idea applies to text that needs several bytes per visible symbol:

text               👋
UTF-8 bytes        F0  9F  91  8B
no useful merges   [F0] [9F] [91] [8B]    4 tokens
learned sequence   [F09F918B]              1 token

Whether the emoji becomes one token depends on the training corpus and vocabulary budget. A common byte sequence is likely to receive a shortcut; a rare sequence may remain split into several tokens. Either way, every byte is still representable.

GPT-style byte BPE. GPT-2 and OpenAI's tiktoken encodings use this byte-based BPE idea. GPT-2 maps raw byte values to visible Unicode symbols while it learns and stores merges, because many raw bytes are not printable. The visible symbols are an internal representation; each learned token ultimately stands for one byte or a sequence of bytes. Modern vocabulary sizes vary by model and include these learned byte sequences plus special tokens.

Why not just use Unicode code points as the vocabulary?

This is the obvious question, and it deserves a real answer. Unicode already numbers every character. Python even hands you the number for free with ord("A") == 65. Why not declare “vocabulary = Unicode” and skip training a tokenizer entirely?

Because it loses on both of the axes a tokenizer is judged on, at the same time.

Problem 1 — far too many tokens

One code point per token means no compression whatsoever. The sentence Tokenization is the first step of every language model. is 9 words and 55 code points, so it costs 55 tokens. A trained subword tokenizer encodes the same sentence in about a dozen. Every token you spend is a position in the context window and a step of attention — and attention cost grows quadratically with sequence length, so a sequence roughly 4–5× longer is far more than 4–5× more expensive.

Problem 2 — a huge and mostly dead vocabulary

A larger vocabulary requires the model to store and score more token IDs. This affects both the beginning and the end of the transformer.

  1. At the beginning: the input embedding table stores one learned vector for every token ID. If the vocabulary has 50,000 tokens, it needs 50,000 rows. Each row has d_model numbers describing that token.
  2. At the end: the output layer compares the transformer's current hidden state with every possible next token. It produces one raw score, called a logit, for each vocabulary entry.
  3. Softmax: converts those logits into probabilities. For example, cat: 0.60, dog: 0.25, and all other tokens sharing the remaining probability.

The input table has shape vocab_size × d_model. The output projection has the opposite shape, d_model × vocab_size, but both contain the same number of values: vocab_size × d_model. Many models use weight tying, which means these two layers share one set of parameters. Models without weight tying pay this parameter cost twice.

Consider a small transformer with d_model = 768. One matrix needs:

number of parameters = vocabulary size × 768

50,257 × 768 = 38,597,376 parameters ≈ 38.6 million

The same calculation shows how quickly the cost grows:

Vocabulary Size Parameters in one matrix What this means
Raw bytes 256 0.2 million Very small vocabulary, but raw text needs many tokens
Example byte-level BPE 50,257 38.6 million Larger vocabulary, but common byte sequences become shorter
Every assigned Unicode character 159,801 122.7 million Large table with many characters a particular model may rarely use

The Unicode table in this example needs about three times as many parameters as the BPE table. The model must also produce and compare 159,801 logits at every output position instead of 50,257. A bigger vocabulary can shorten token sequences, but it also makes the embedding and output layers wider. Tokenizer design is a balance between those two costs.

Worse, most of that table is dead weight. An English corpus touches maybe a hundred distinct code points; the remaining ~159,700 rows receive almost no gradient and stay essentially random. You would be paying for capacity you never train.

Problem 3 — it is still not closed

You might accept the cost for the guarantee that nothing is ever unknown. You do not get that guarantee. A character-level vocabulary is built from the characters in your training corpus, so any character outside it is still [UNK]. And Unicode itself keeps growing — version 17.0 added 4,803 characters and 4 new scripts. A vocabulary pinned to a Unicode version goes stale.

The verdict. Unicode-character tokenization gives you long sequences and a large vocabulary and still no coverage guarantee — the worst corner of the tradeoff. Bytes fix the second and third problems outright, and BPE on top of bytes fixes the first.

Character level vs byte level, side by side

Character level (Unicode) Byte level (UTF-8)
Atomic unit One Unicode code point One byte, 0–255
Base vocabulary size Open-ended — up to ~160,000 and growing every release Fixed at exactly 256, forever
Unseen input [UNK] — the spelling is lost Impossible; every byte is already in the vocabulary
Reuse across scripts None. é, e, and আ are unrelated IDs with no shared structure Heavy. The same 256 atoms are reused by every script on earth
Cost of one character Always 1 token 1–4 tokens before merges; usually back under 1 after BPE
Rows that stay untrained The vast majority, on any single-language corpus Essentially none — all 256 bytes get used

The byte column’s “1–4 tokens per character” looks like a loss, and on raw bytes it is: বাংলা ভাষা is 10 code points but 28 UTF-8 bytes. But those bytes are only the starting alphabet. BPE then merges the recurring byte sequences of Bengali into single pieces, and the sequence collapses back down. You get the coverage of bytes with the length of subwords.

Byte fallback: the hybrid most models actually ship

You do not have to choose one or the other. Byte fallback keeps a normal subword vocabulary for everything common, and drops to raw bytes only for things that vocabulary cannot express:

  • Text that matches a learned piece → encoded as that piece, as usual.
  • Anything else — a rare script, a brand-new emoji, a corrupted byte — is emitted as individual byte tokens such as <0xF0> <0x9F> <0x98> <0x8A>.

The result is that [UNK] never appears and decoding is always exact, at the price of a few extra tokens in rare cases. This is SentencePiece’s byte_fallback option, and it is what the Llama tokenizers use. It is the practical answer to “what if something cannot be represented?” — the fallback is the raw byte.

UTF-8 inspector

Bytes

Grapheme Code point(s) CP # UTF-8 bytes Byte #

Leave when: you can read one row of the inspector and separate grapheme / code point / byte counts.

Vocabulary size, context length, and sequence length

Three numbers that beginners routinely merge into one. They are independent, they are set at different times, and confusing them produces confident nonsense about what a model can do.

Does “1B” mean one billion vocabulary tokens?

No. In a model name such as Llama 3.2 1B, the 1B means the neural network has roughly one billion learned parameters. Likewise, 7B means about seven billion parameters and 70B means about seventy billion. The letter B means billion; it does not describe the tokenizer vocabulary.

A parameter is a number adjusted while the model is trained. Together, billions of these numbers let the network learn language patterns and store useful information. Parameters are spread across the embedding layer, attention layers, feed-forward layers, and other parts of the model.

Number What it measures Example meaning
1B, 7B, 70B Total model parameters Approximately 1, 7, or 70 billion learned numbers
vocab_size: 128256 Tokenizer vocabulary 128,256 different token IDs are available
context_length: 131072 Context capacity Up to 131,072 token positions can fit in one request

These numbers do not have to grow together. A two-billion-parameter model can use a vocabulary of 32,000 tokens, 128,000 tokens, or another size chosen by its designers. To find the real vocabulary size, check the model configuration or tokenizer files for a field such as vocab_size.

How vocabulary contributes to parameters. If a model has a vocabulary of V tokens and each token embedding contains D numbers, its embedding table contains V × D parameters. A larger vocabulary therefore increases the model's parameter count, but it does not determine the model's name. Most parameters are usually in the transformer blocks, although embeddings can take a noticeable share of a small model with a large vocabulary.
model name:       7B              → total learned parameters
tokenizer config: vocab_size      → number of available token IDs
model config:     context_length  → maximum token positions per request

What a vocabulary size of 50,257 actually means

GPT-2’s vocabulary size is 50,257. That number is not arbitrary — it decomposes exactly:

Component Count What it is
Base byte tokens 256 Every possible UTF-8 byte — the atoms nothing can fall through
Learned merges 50,000 The merge budget BPE was given during training
Special token 1 <|endoftext|>
Total 50,257
Read this carefully. A vocabulary size of 50,257 means there are 50,257 unique token pieces — not 50,257 words. Most entries are not words at all: single bytes, punctuation, whitespace runs, and word fragments like ing, tion, or the (with its leading space attached).

And because pieces combine, the set of strings the tokenizer can represent is not capped at 50,257. A word that is not in the vocabulary is simply split into smaller pieces that are, then reassembled on decode:

unhappiness       →  un + happiness              2 pieces
tokenization      →  token + ization             2 pieces
Kubernetes        →  Kub + ernet + es            3 pieces
zyxwvu            →  z + y + x + w + v + u       6 pieces (worst case: one per byte)

Sequences of pieces multiply. With 50,257 pieces, even two-piece combinations already exceed two billion distinct strings, and there is no limit on sequence length. Therefore BPE can represent vastly more than 50,257 words — in fact every string that can be written in UTF-8, including words that will be invented next year. The vocabulary size caps how efficiently text is encoded, never whether it can be encoded.

That is the whole reason bigger vocabularies keep shipping: more pieces means fewer pieces per sentence, not more words covered. Coverage was already total.

Tokenizer Algorithm Vocabulary size
BERT base uncased WordPiece 30,522
Llama 2 SentencePiece BPE 32,000
GPT-2 (r50k_base) Byte-level BPE 50,257
GPT-4 (cl100k_base) Byte-level BPE 100,277
Llama 3 Byte-level BPE 128,256
GPT-4o (o200k_base) Byte-level BPE 200,019

Vocabulary size is not context length

These answer two completely different questions:

Vocab size
How many different tokens exist. The size of the dictionary. Fixed when the tokenizer is trained, and frozen for the life of the model — changing it means retraining the embedding table.
Context length
How many tokens the model can process at once. The number of positions in the window. Set by the model architecture and training, and has nothing to do with how many distinct tokens exist.
Sequence length
How many tokens your particular input came out as. Varies per input, and must fit inside the context length.

GPT-2 makes the independence obvious: vocabulary 50,257, context length 1,024. The dictionary has fifty thousand entries; the model can read a thousand tokens at a time. Neither number constrains the other.

Analogy. Vocabulary size is how many words are in your dictionary. Context length is how many words you can hold in your head at once. Learning more words does not enlarge your working memory — but it does let you say the same thing in fewer words, so more meaning fits.

That last clause is the one real connection between them. A larger vocabulary does not change the context length, but it lowers the token count of the same text, so more text fits in the same window. This is why o200k_base encodes non-English text in noticeably fewer tokens than cl100k_base despite the window being a property of the model, not the tokenizer.

See the tradeoff

Same sentence, four granularities. The bars are token counts. This is the tradeoff every tokenizer paper is negotiating.

Length comparison

Counts

Why don’t modern models use word-level tokenization if it uses fewer tokens?

Word-level tokenization looks efficient in this chart because every word in the sample sentence is familiar. In real text, that advantage quickly breaks down.

  1. An unknown word loses its identity. A word vocabulary is frozen before model training. If COVID-19, useState, a username, a rare name, or a misspelling is missing, the tokenizer returns [UNK]. The model does not receive the word's spelling, so it cannot see which known parts it contains. A subword tokenizer can preserve that information using smaller pieces; byte-based subwords can always fall back to bytes.
  2. Related word forms become unrelated IDs. A word tokenizer stores play, plays, played, and playing as four independent entries. Subwords can reuse pieces such as play, ed, and ing. This matters even more for languages with many word forms and productive compounds, including Turkish, Finnish, Arabic, and German.
  3. Better coverage requires an enormous vocabulary. Adding more whole words reduces [UNK], but every entry needs an embedding and an output score. A vocabulary with hundreds of thousands of words makes those layers larger and makes next-token scoring more expensive. Rare entries may still appear too few times to learn useful representations.
  4. Multilingual text does not fit neatly into one word list. A shared word vocabulary would need entries for words from every supported language, plus mixed-language text, emoji, and code. Subword and byte-based tokenizers reuse a smaller set of pieces across all of them.
  5. The token saving is smaller than it first appears. Keeping a common word as one token instead of two saves one position. That modest saving is not worth losing every rare or newly invented word. Subwords accept slightly longer sequences for common text in exchange for reliable coverage of the long tail.
training vocabulary contains:  play, football, the

word level
playing      → [UNK]
footballer   → [UNK]
useState     → [UNK]

subword level
playing      → play + ing
footballer   → football + er
useState     → use + State
The real tradeoff. Word-level tokenization gives the shortest sequence for known words, but fails on unknown words and needs a vocabulary that grows with every language and word form. Subwords use a few more token positions, but keep the vocabulary manageable, reuse structure, and preserve unfamiliar text. That is why modern language models generally choose subword or byte-based subword tokenization.

Leave when: you can say which bar is not a production tokenizer, and explain to someone else why a 50,257-piece vocabulary can represent far more than 50,257 words while saying nothing at all about context length.

Visuals

Five printable schematics (Token Lab cyan on light paper). Open full-page for projectors — better than mermaid fences when Jupyter does not render diagrams.

01 · Tokenizer pipeline Raw → normalize → pre-tokenize → model → IDs → decode
02 · Granularity tradeoff Word / subword / character / byte layers
03 · BPE train loop Count pairs, pick max frequency, merge, repeat
04 · BPE vs WordPiece Frequency vs likelihood score, then shared glue
05 · Detokenize marks </w>, ##, and ▁ → string

Gallery: diagrams/

Glossary

Unicode
The universal standard that assigns a unique code point to every character: A → U+0041, 😊 → U+1F60A. It defines which characters exist, not how they are stored.
Code point
One numbered entry in the Unicode catalogue, written U+XXXX in hex. 159,801 of the 1,114,112 possible code points are assigned as of Unicode 17.0.
UTF-8
The encoding that turns Unicode code points into bytes for storage and transmission. ASCII characters take 1 byte; others take 2–4. Unicode defines the characters, UTF-8 defines the bytes.
Byte
A value 0–255. There are exactly 256 of them and that number never grows — the appeal of byte-level tokenizers.
Grapheme
What a reader calls “one character” on screen. May be several code points (a letter plus a combining accent, or an emoji sequence).
Byte fallback
Emitting raw byte tokens such as <0xF0> for anything the subword vocabulary cannot express. Removes [UNK] entirely. Used by the Llama tokenizers via SentencePiece byte_fallback.
Vocab size
How many distinct token pieces exist. GPT-2: 50,257 = 256 bytes + 50,000 merges + 1 special token. Pieces, not words — they combine, so coverage is unlimited.
Context length
How many tokens the model can process at once. A property of the model, entirely separate from vocab size.
Logits
The raw scores the final layer produces — one per vocabulary entry, before softmax. A bigger vocabulary means a wider logit vector at every position.
Embedding
The vector a token ID is looked up as. The lookup table is vocab_size × d_model, which is why vocab size costs parameters.
Greedy
Takes the best option available at the current step, with no lookahead and no backtracking. BPE merge selection is greedy.
OOV / [UNK]
Out of vocabulary. English: unknown word. বাংলা: অজানা শব্দ — the model sees a hole.
Type
A distinct vocabulary entry (the string), not each occurrence in a sentence.
Merge
BPE / WordPiece training step that glues two adjacent symbols into one new piece.
##
WordPiece continuation mark. ##est continues a word; it does not start one.
</w>
BPE end-of-word mark. Distinguishes the st in star from the end of widest.
▁ (U+2581)
SentencePiece space mark — not a plain underscore. Decoding turns each ▁ into a space.
NFKC
Unicode normalization form. Makes compatibility-equivalent characters look the same before training.

Next: four tokenization playbooks

The labs above are intuition. The notebooks build the machines from scratch on the Sennrich toy (or a tiny Viterbi demo), then repeat the same idea on data/tiny_corpus.txt with the current PyPI libraries via uv. The fourth playbook compares traditional and byte-level BPE directly.

01_bpe.ipynb — Byte Pair Encoding Frequent pair merges, HuggingFace BpeTrainer, then tiktoken as production BPE. Open BPE Explained playbook in Colab
02_wordpiece.ipynb — WordPiece Likelihood score instead of raw frequency, ## continuation, BERT-style trainer.
03_sentencepiece.ipynb — SentencePiece No whitespace pre-tokenizer, ▁ space mark, Unigram + BPE modes, sentencepiece 0.2.2 API.
04 — Byte-level BPE vs traditional BPE Same merge algorithm, different base symbols: observed Unicode characters vs a universal 256-byte vocabulary. Open byte-level BPE comparison playbook in Colab

From the repo root:

uv sync
uv run jupyter lab notebooks

Repo: github.com/gaurav36/tokenization-explanation · Live site: gaurav36.github.io/tokenization-explainer

To do after Notebook 03

Tokenize the same real prompt with three model tokenizers and compare the results:

  1. BERT: use its WordPiece tokenizer and look for continuation pieces marked with ##.
  2. GPT-4o: use its OpenAI tiktoken encoding and inspect how spaces, punctuation, and UTF-8 text are grouped.
  3. T5: use its SentencePiece tokenizer and look for the ▁ symbol that represents a word-starting space.

For each tokenizer, record the token pieces and token count. Then explain why the counts differ even though the input text is identical. Use each model's real tokenizer: tiktoken is appropriate for GPT-4o, but BERT and T5 require their own tokenizer implementations.