Think of tokenization as a translator. You write normal text, such
as Hello world, and the tokenizer turns it into numbers like
[101, 7592, 2088, 102]. The model only ever sees those numbers.
Those numbers come from a fixed vocabulary — simply a list of all
the pieces the tokenizer knows. Here, 101 and 102 mark the
start and end of the sentence, while 7592 and 2088 stand
for the pieces in between. When the model finishes, the tokenizer runs in reverse and
turns the numbers back into readable text.
Tokenizers such as BPE, WordPiece, and
SentencePiece are the recipes for choosing those pieces — and for
reversing the process afterwards.
Never seen this before? Start here.
A language model cannot read. It can only do arithmetic. So every piece of text has
to be turned into numbers before the model sees it, and turned back into text
afterwards. Tokenization is that conversion.
The only real question in this whole topic is: how big is one piece?
A whole word? Half a word? One letter? One byte? Every answer trades something away,
and this page lets you feel each trade in a live playground instead of taking it on
trust.
A neural network works with numbers, not with words or sentences. So before a
transformer can process text and predict what comes next, the text must first be
converted into a sequence of token IDs.
For example, a BERT-style tokenizer might turn Hello world into
[101, 7592, 2088, 102]. Here, 101 and 102 are
special tokens that mark the beginning and end of the sequence, while
7592 and 2088 represent the pieces of text in between.
The transformer does not receive the original words. It receives these IDs, which are
then turned into numerical vectors called embeddings. After the model
produces its output, the tokenizer is used again to turn token IDs back into readable
text.
At a high level — the full round trip
TextHello world
Tokenizerencode
Tokenshello · world
Token IDs7592 · 2088
Embeddingsvectors
Transformerthe model
Output IDspredicted
Tokenizerdecode
Textreadable again
what a human reads the tokenizer, used twice numbers only the model
The one takeaway. The tokenizer is the bridge between
human-readable text and the
numbers a neural network can process. It is the only part of the stack
that speaks both languages, which is why it appears at both ends of the diagram.
Example: if two tokenizers both see Hello world, they may agree on the
idea of “hello” and “world” but still assign different IDs. The number is just a label
inside that tokenizer’s own vocabulary.
The Tokenization Pipeline
Tokenization is not just "splitting a sentence into words." A tokenizer usually
performs several steps, and each step has a specific job.
Think of the process as a pipeline. Each stage prepares the text for the next one. If
a stage is skipped or done badly, you can get wrong lengths, unknown tokens, or broken
multilingual text.
Example: "Hello, I'm learning AI!" may first stay as raw text, then be
normalized, then split into pieces like Hello, ,,
I, ', m, learning, AI,
! before those pieces are mapped to IDs.
Here is the full path from text to model input and back again:
The granularity of step 4 is the main design choice. Each level answers one
question: what is one token?
Example: the word lowest can be a single word token, two subword pieces
like low + est, six characters, or six bytes before any
merges happen. The answer depends on the tokenizer you choose.
Word
One token ≈ one whitespace word. Tiny sequences, huge vocabulary, and any new word
becomes unknown. Fine for small closed domains. Fragile for open web text.
Example: I love football becomes three tokens, but
ChatGPT may become [UNK] if the word was never trained.
Subword
One token is a frequent chunk: token + ##ization, or
low + er. This is how BERT, GPT, T5, and Llama actually
work. Three training recipes live in the algorithm labs below.
Example: playing might become play + ing, so
the model can reuse pieces it already knows instead of learning every word from
scratch.
Character
One token is a Unicode code point (sometimes a grapheme cluster). Vocabulary stays
small. Sequences get long. Spelling and morphology become the model’s problem.
Example: hello becomes h, e, l,
l, o. A short word looks harmless, but a whole sentence
gets much longer very quickly.
Byte
One token is a UTF-8 byte (0–255). Nothing is unknown — every Unicode string is
some byte sequence. GPT-2-style BPE starts here, then merges bytes into longer
pieces.
Example: A is 1 byte, é is 2 bytes, and 👋 is
4 bytes before merges. The model always has a way to spell the text.
Rule of thumb. Coarser tokens → shorter sequences, bigger vocab,
more out-of-vocabulary pain. Finer tokens → longer sequences, smaller vocab, more
compute per sentence. Subword is the industry compromise.
Example: a word tokenizer may use one token for house, a subword tokenizer
may use house or hous + e, a character tokenizer
uses five tokens, and a byte tokenizer starts with the UTF-8 bytes underneath it.
What a vocabulary actually is
A vocabulary is a two-way map: piece ↔ integer. Special pieces almost always sit at
the front: padding, unknown, start, end, mask. Training a tokenizer means
choosing the piece list so that a large corpus compresses well. Using a
tokenizer means looking pieces up — you do not invent new IDs at inference
time.
Token
One atomic piece after splitting. Might be a word, a subword, a character, or a byte.
Type / ID
The integer that names that piece inside the vocabulary. The model embeds this ID.
OOV / UNK
Out of vocabulary: a piece the vocab never listed. Word-level tokenizers hit this constantly.
Detokenize
Glue pieces back into a string. Spaces, ##, and the SentencePiece ▁ all exist to make this reversible.
Leave when: you can name the six pipeline stages and say which one
chooses word vs subword vs byte.
Unicode and UTF-8
Before a tokenizer ever runs, your text is already numbers. Two different
standards did that job, and beginners almost always blur them together. Separating
them makes everything after this page easier.
Unicode: the catalogue of characters
Unicode is a universal standard that assigns a unique
code point — a number — to every character. It is a giant numbered
catalogue that everyone in the world agrees on:
A → U+0041 (decimal 65)
é → U+00E9 (decimal 233)
আ → U+0986 (decimal 2438)
😊 → U+1F60A (decimal 128522)
The U+ prefix just means “this is a Unicode code point, written in hex.”
Because every language, symbol, and emoji has its own agreed number, a document written
in Tokyo opens correctly in Zürich. That is the whole point of Unicode: consistent
identity for characters.
Example: the letter A is always U+0041, whether it appears in
a sentence, a filename, or a code comment. Unicode gives the character its identity
before any tokenizer decides what to do with it.
Crucially, Unicode says nothing about how those numbers are stored in
a file or sent over a network. It only hands out the numbers.
UTF-8: the rule for storing those numbers as bytes
UTF-8 is an encoding: it converts Unicode code points into
bytes so text can be stored on disk or transmitted. It uses a variable number of bytes
per character:
ASCII characters (U+0000–U+007F) use 1 byte
Everything else uses 2, 3, or 4 bytes
The relationship in one line.
Unicode defines which characters exist and what number each one has → UTF-8
defines how that number is written down as bytes. Unicode is the catalogue;
UTF-8 is the packing rule.
Example: A is one byte in UTF-8, but 😊 is four bytes. That is
why byte-level tokenizers see more pieces for emoji and many non-Latin scripts until
later merges compress them again.
The same four characters, all the way down
Read this table left to right and you have followed one character from human symbol,
to Unicode number, to the actual bytes on disk:
Character
Unicode name
Code point
UTF-8 bytes (hex)
Bytes
A
LATIN CAPITAL LETTER A
U+0041
41
1
é
LATIN SMALL LETTER E WITH ACUTE
U+00E9
C3 A9
2
আ
BENGALI LETTER A
U+0986
E0 A6 86
3
日
CJK UNIFIED IDEOGRAPH-65E5
U+65E5
E6 97 A5
3
😊
SMILING FACE WITH SMILING EYES
U+1F60A
F0 9F 98 8A
4
Notice that A costs one byte and 😊 costs four. That single
fact explains the byte column in the
four-way lab below, and it is why the same sentence can be cheap in
English and expensive in Bengali or Japanese.
Example: café looks short on screen, but the accent means the code-point
and byte counts are different from plain cafe. That is why the inspector
below separates graphemes, code points, and bytes.
Build your own rows in the
UTF-8 inspector — type anything and watch graphemes, code points,
and bytes diverge.
Three different ways to count “length”
Beginners lose hours to this. The string café 👋 has three defensible
lengths, and each one is correct for a different question:
Example: a browser may show 👋 as one visible character, but a string API
can count code points differently, and the byte-level tokenizer always counts the UTF-8
bytes underneath it.
Grapheme
What a human calls “one character” on screen. é is one grapheme even
when it is built from e + a combining accent.
Code point
One entry in the Unicode catalogue. é can be
one code point (U+00E9) or two
(U+0065 + U+0301) — which is exactly what
normalization (NFKC in the pipeline) is for.
Byte
One UTF-8 byte, 0–255. This is what is actually stored, and what byte-level
tokenizers consume.
How big is Unicode, really?
1,114,112code points in the Unicode codespace (U+0000–U+10FFFF)
159,801characters actually assigned as of Unicode 17.0 (Sept 2025)
172scripts (writing systems) supported — serving hundreds of languages
256possible values of one UTF-8 byte — and that never grows
Note the wording: 172 scripts, not 172 languages. One script such as Latin or
Devanagari is shared by many languages. Also note that the assigned-character count
grows with every Unicode release — the 256 byte values never do. Hold on to that
contrast; section 10 turns it into a design decision.
Why UTF-8 won. It is backward compatible with ASCII (English text is
byte-for-byte identical to a 1960s ASCII file), it has no byte-order ambiguity, it can
encode every Unicode character, and it is self-synchronizing — you can jump into the
middle of a stream and find the next character boundary. Python and JavaScript hold
strings differently in memory (Python uses code points, JavaScript uses UTF-16 units),
but files, HTTP, and model tokenizers overwhelmingly speak UTF-8.
Leave when: you can say in one sentence what Unicode gives you, what
UTF-8 gives you, and why 😊 is four bytes but A is one.
Four-way lab
Type anything. The same string is split four ways at once. Watch how token
count explodes as units get smaller, and how punctuation, emoji, and non-English
scripts behave.
Example: try café 👋 or বাংলা 日本語. The word view stays short,
the character view grows, and the byte view grows fastest before merges have a chance to
compress repeated patterns.
Live splitter
Word
0
Naive subword
0
Character
0
UTF-8 byte
0
“Naive subword” here is a teaching trick: 3-character chunks with WordPiece-style
##. Real BPE / WordPiece / SentencePiece learn those chunks from data —
open Algorithm labs.
Leave when: you can explain why the byte column is longer than the
character column for 👋.
Word level
The oldest idea: split on whitespace (and usually punctuation). Each distinct word
string gets its own ID. “cat” and “cats” are unrelated. “Tokenization” is a brand-new
type even if the model already knows “token”.
Example: I love football becomes three IDs in a word tokenizer, but
I love ChatGPT can fall apart if ChatGPT was never seen in
training. That is the OOV problem in one sentence.
That is the out-of-vocabulary (OOV) problem. At training time you only
store the words that appeared often enough. At test time a name, a typo, or a new
product word becomes [UNK]. The model then sees a hole, not the spelling.
The OOV trap
Lookup
Word IDs after lookup
Try replacing dog with cats, a name, or a Bengali word. Word-level
tokenization has no way to reuse letters it already knows. That single failure is why
modern LLMs are not word tokenizers.
Leave when: you can say why cat and cats are
unrelated IDs under word-level tokenization.
Subword level
Subword tokenization keeps frequent words intact and splits rare words into reusable
pieces. unhappiness might become un + happiness,
or un + happy + ness, depending on the algorithm
and the corpus.
Example: playing can stay whole if it is common, but a rarer word like
unhappiness is more useful when split into pieces that other words can
reuse. That is why subword tokenization feels like a compromise: it keeps common words
short and rare words possible.
Same corpus, three scoring rules. Play all three in the lab, then train the real libraries in Jupyter.
BPE
Merge the most frequent adjacent pair, again and again. GPT family, many Llama tokenizers.
WordPiece
Merge the pair that most improves a likelihood score, not raw count. BERT, DistilBERT. Continuation mark ##.
SentencePiece
Treat the raw string as the unit (spaces become the character ▁). Unigram LM or BPE. T5, ALBERT, many multilingual models.
Misconception to kill. BPE pieces are not guaranteed morphemes.
est appears because the pair was frequent, not because English has a
comparative suffix. Probe star vs widest with and without
</w> in mind.
Example: a tokenizer might learn low + est from the corpus,
then reuse those same pieces in lowest and widest even though
the words are not built the same way in grammar textbooks. The pieces are learned from
frequency, not from a dictionary of linguistics rules.
Leave when: you can name the scoring rule for each of the three
algorithms in one sentence.
Before learning BPE, it helps to see where the main tokenizer methods are used.
Vocabulary size means the number of token IDs the tokenizer can
produce. It does not mean the number of words or ideas the model understands.
Model version
Tokenizer method
Starts with
Vocabulary size
GPT-2
Byte-level BPE
UTF-8 bytes
50,257
RoBERTa
Byte-level BPE
UTF-8 bytes
50,265
BERT base, uncased
WordPiece
Characters
30,522
ALBERT base v2
SentencePiece Unigram
Characters
30,000
Llama 2
SentencePiece BPE
Characters and bytes
32,000
Llama 3
Tiktoken-based BPE
UTF-8 bytes
128,256
DeepSeek-V3
Byte-level BPE
UTF-8 bytes
129,280
Qwen2.5
Byte-level BPE
UTF-8 bytes
152,064
Kimi K2
Byte-level BPE
UTF-8 bytes
163,840
Llama 3.3 70B on Groq
Tiktoken-based BPE
UTF-8 bytes
128,256
Always check the exact model version. Models in the same family can
use different tokenizers. Vocabulary counts can also include special tokens such as
[CLS], [SEP], or chat-control tokens, so sources sometimes
report slightly different totals.
Why is Groq listed differently? Groq runs models from companies such
as Meta, Google, and OpenAI on its inference platform. It does not give every hosted
model one “Groq tokenizer.” A Llama model served by Groq still uses its Llama
tokenizer, while a different hosted model uses its own tokenizer.
How to train your own tokenizer
You can train a tokenizer without training a language model. The tokenizer learns
useful pieces from a collection of text called a training corpus.
Collect representative text. Include the languages, code, names,
numbers, and writing styles that your model will receive.
Choose a method. Use BPE, WordPiece, or SentencePiece Unigram,
depending on the behaviour you want.
Choose a vocabulary size and special tokens. For example, request
32,000 token IDs and reserve tokens such as [UNK], [BOS],
and [EOS] if your model needs them.
Run the tokenizer trainer. It reads the corpus, counts patterns,
and builds the vocabulary and rules. No neural-network training is involved yet.
Test, save, and freeze it. Check common text, rare words, every
target language, emoji, and code. Once model training begins, keep the token IDs
fixed so their learned embeddings do not change meaning.
Try it in this project. The three Jupyter
notebooks train BPE, WordPiece, and SentencePiece tokenizers on a small corpus, so
you can inspect the learned vocabulary before using a large dataset.
How BPE works, step by step
Byte Pair Encoding is the algorithm behind the GPT family and most Llama tokenizers.
It is also genuinely simple — simple enough to do by hand, which is what this section
does before you drive it in the lab.
The algorithm
BPE is a greedy algorithm that builds a vocabulary by repeatedly
merging the most frequent pair of adjacent tokens. “Greedy” means it takes the best
pair available right now, with no lookahead and no going back to reconsider an
earlier merge.
Prepare the corpus. Split the text into words and count how often
each distinct word occurs. Mark word boundaries — this lab uses an end-of-word
symbol </w> — so merges can never run across two words.
Start with individual characters as tokens. Every word is spelled
out as a sequence of single characters. Those characters are the starting
vocabulary. (Byte-level BPE starts from the 256 bytes instead — same
algorithm, different atoms.)
Count adjacent pairs. Across the whole corpus, count how often each
pair of neighbouring tokens occurs, weighted by how often its word occurs.
Merge the most frequent pair into one new token, everywhere it
appears.
Record the merge and add the new token to the vocabulary. The merge
list is ordered — that order matters later.
Repeat steps 3–5 until the vocabulary reaches its target size (or,
equivalently, until you have performed a chosen number of merges). This is the stop
condition.
Why BPE must stop merging
First, a merge means joining two neighbouring tokens and treating
them as one new token. Suppose BPE starts with the word low split into
three character tokens:
start l + o + w 3 tokens
merge l + o lo + w 2 tokens
merge lo + w low 1 token
This is what “the word shrinks” means: the text does not lose letters, but the number
of tokens becomes smaller. If BPE continued merging every possible pair, each word
in the training corpus would eventually become one large token. For example,
lowest would no longer be reusable pieces such as low +
est; it would become the single token lowest. The vocabulary
would then behave too much like a word-level vocabulary, with a separate large token
for each training word.
Real BPE therefore stops at a limit chosen before training. That
limit may be a target vocabulary size or a fixed number of merges. Think of it as a
spending budget: BPE is allowed to create only a certain number of new tokens. GPT-2,
for example, was given 50,000 merges. This keeps useful common pieces while avoiding
a separate token for every complete word.
A second stopping rule can ignore very rare pairs. If a pair appears only once,
creating a permanent vocabulary entry for it usually saves very little. This lab
merges pairs with frequency 2 or more and stops when the best remaining pair has
frequency below 2.
What happens after BPE stops? Training is finished. BPE saves the
new vocabulary pieces and the merge rules in the exact order they were learned.
These are then frozen: when the tokenizer receives new text, it does not count pairs
or learn new merges. It simply starts from the smallest units and replays the saved
rules that apply.
For example, imagine training stops after learning these four rules:
01 l + o → lo
02 lo + w → low
03 e + s → es
04 es + t → est
The saved vocabulary now includes pieces such as lo, low,
es, and est. When BPE later sees lowest, it
replays those rules and produces low + est. If it sees
lower, it can produce low + e + r
because no saved rule joins the final e and r. Each resulting
piece is then looked up in the vocabulary and replaced with its token ID for the
model.
Training chooses the pieces; encoding uses them. After the stopping
point, the tokenizer no longer learns. It applies the frozen merge list to every new
input and falls back to smaller known units whenever no larger saved merge applies.
For example, with the four saved rules above:
lowest → low + est
lower → low + e + r
lovely → lo + v + e + l + y
lowest can use two large learned pieces. lower can use
low, but its remaining letters stay separate because no
e + r rule was learned. lovely can use only
l + o → lo, so everything else falls back to individual characters.
The tokenizer does not create a new lovely token while encoding it.
What is the disadvantage? Frozen rules make token IDs stable, but
they also preserve the biases of the training corpus. Text similar to the training
data gets large, efficient pieces; unfamiliar words, languages, or technical terms
may be split into many small pieces.
Longer sequences:lowest uses 2 token positions,
while lovely uses 5 in this example.
More computation: the transformer must process every token, so
more pieces require more work and consume more of the context window.
Uneven efficiency: a language or subject that was rare in the
training corpus may need far more tokens than common English text.
Difficult to update: adding new vocabulary pieces later changes
the token-ID system and normally requires updating or retraining the model's
embedding and output layers.
The text is still representable, so fallback is better than producing
[UNK]. The disadvantage is mainly inefficiency, not
loss of the original spelling.
BPE must know where one word ends and the next word begins. In the cat,
it may join letters inside the, such as t + h → th, or
inside cat, such as c + a → ca. But it must never join
e + c, because e belongs to the and
c belongs to cat.
The marker </w> means “this word ends here.” The sentence can be
represented as t h e </w> and c a t </w>, so BPE knows
not to merge across the boundary.
The starting pieces depend on the BPE version. Character-level BPE starts with
letters such as c, a, and t. Byte-level BPE
starts with byte values from 0 to 255. After that, both versions follow the same
basic process: count neighbouring pairs, merge common pairs, and stop at the chosen
limit.
Worked example: the classic four-word corpus
This is the corpus loaded in the lab below — four word types with their counts:
The × number tells you how many times that complete word appears
in the training corpus. For example, low ×5 means the corpus
contains five copies of low, while lower ×2 means it contains
two copies of lower. The word is shown only once here to keep the table
compact. spelled shows the starting character tokens, and
</w> marks the end of the word.
low ×5 spelled l o w </w>
lower ×2 spelled l o w e r </w>
newest ×3 spelled n e w e s t </w>
widest ×2 spelled w i d e s t </w>
Step 3 counts every adjacent pair. The pair l+o appears in
low (5 times) and in lower (2 times) — 7 occurrences, the
highest in the corpus. So it wins the first merge. Repeat, and BPE produces this exact
ordered merge list:
In the merge list, freq means the total number of times the
adjacent pair appears across the corpus at that step. Word counts are included
in that total. For example, l + o appears once inside each copy of
low and lower, so its frequency is
5 + 2 = 7. After it is merged into lo, the pair
lo + w also appears seven times. The frequency is recalculated after every
merge because the available token pairs have changed.
01 l + o → lo freq 7
02 lo + w → low freq 7
03 low + </w> → low</w> freq 5
04 e + s → es freq 5
05 es + t → est freq 5
06 est + </w> → est</w> freq 5
07 n + e → ne freq 3
08 ne + w → new freq 3
09 new + est</w> → newest</w> freq 3
10 low + e → lowe freq 2
11 lowe + r → lower freq 2
12 lower + </w> → lower</w> freq 2
13 w + i → wi freq 2
14 wi + d → wid freq 2
15 wid + est</w> → widest</w> freq 2
After merge 15 no pair is left with frequency ≥ 2, so training stops on the guard
rather than the budget. Reproduce every line of this by dragging the slider in the
BPE lab — the merge log there is generated by the same
algorithm.
Encoding a word BPE has never seen
This is the payoff. The word lowest is not in the
training corpus. To encode it, BPE spells it out in characters and then replays the
learned merges in the order they were learned:
start l o w e s t </w>
after 01 lo w e s t </w> (l+o)
after 02 low e s t </w> (lo+w)
after 04 low es t </w> (e+s)
after 05 low est </w> (es+t)
after 06 low est</w> (est+</w>)
result ["low", "est</w>"] → 2 tokens, decodes back to "lowest"
Merge 03 (low+</w>) does not apply, because in
lowest the piece low is not followed by the end of the word.
That is precisely why the </w> mark exists: it lets BPE distinguish
“low as a complete word” from “low as the start of a longer word.”
Advantage and tradeoff. The advantage is total coverage:
BPE can represent a word it has never seen by using smaller known pieces, so byte-level
BPE never needs [UNK]. The disadvantage is inefficiency:
an uncommon word may need many tokens, which uses more context space and computation.
In the worst case, the tokenizer emits one character or one byte at a time.
The strawberry and blueberry problem
People see the shared ending berry immediately. BPE does not begin with
that meaning. It only learns whichever neighbouring pieces were frequent in its
training corpus. One tokenizer might produce:
Another tokenizer, trained on different text or with a smaller vocabulary, might
split the same words less neatly:
strawberry → str + aw + ber + ry 4 tokens
blueberry → blue + b + err + y 4 tokens
Both tokenizations are valid because every piece joins back into the original word.
The second version is less efficient and does not expose berry as one
reusable piece. It can also make questions such as “How many rs are in
strawberry?” harder: the model receives token IDs for chunks, not a tidy
list of individual letters. Tokenization contributes to that difficulty, although it
is not the only reason a language model may answer a spelling question incorrectly.
The exact split depends on the tokenizer, so always inspect the real tokens instead of
assuming these example splits.
What BPE is not
Not a morphological analyser.est emerged because that
pair was frequent, not because English has a superlative suffix. On a different
corpus you get different pieces, and plenty of them straddle real morpheme
boundaries.
Not optimal. Greedy means locally best at each step. There is no
guarantee that this merge sequence is the best possible vocabulary of that size —
only that it is cheap to compute and works well in practice.
Not the same at train and encode time. Training discovers
the ordered merge list from counts. Encoding just replays that fixed list.
No new tokens are ever invented at inference time.
Leave when: you can state BPE’s stop condition correctly, and explain
why running the merge loop “until no more pairs appear” would defeat the purpose.
This lab lets you compare BPE, WordPiece, and SentencePiece on the same
corpus (the text used to learn token pieces). Because all three
trainers receive the same text, you can see how their different learning rules create
different vocabularies and word splits.
Choose an algorithm, move the slider to control how much it learns, and enter a
probe word to see how the trained tokenizer splits new text. The
browser versions are simplified so you can inspect every step. The Jupyter notebooks
later in the lesson repeat the experiments with the production libraries HuggingFace
tokenizers and sentencepiece 0.2.2.
Subword trainers
Sennrich toy = four word types (low / lower /
newest / widest) so WordPiece can form
low + ##est. Course corpus matches
data/tiny_corpus.txt used by HuggingFace / SentencePiece cells in the
notebooks — merges will differ. That is expected.
Slider 0–20 = number of merge steps
Encoded pieces
0
Same probe, three encodings
Frozen to the corpus, slider, and probe above. Switch SentencePiece Unigram/BPE with
the mode buttons — the third column follows.
Default corpus is the Sennrich toy set so WordPiece can form low +
##est. Paste the course corpus and watch WordPiece prefer rare pairs
like qu — that is the likelihood score doing its job, and why production
models train on huge text.
Ready to move on when: you can explain why BPE, WordPiece, and
SentencePiece split the same word into different pieces. The tokenizers agree on the
word; they just disagree on where to split it.
Character level
A character-level tokenizer usually gives each Unicode code point its own
token ID. Its vocabulary contains the distinct characters found in the training text.
An English-focused vocabulary may contain around one hundred letters, digits,
punctuation marks, and special symbols. A multilingual vocabulary may need thousands.
However, one visible symbol is not always one code point:
JavaScript can count one emoji as two. JavaScript's
"👋".length is 2 because length counts UTF-16
code units. Use Array.from(text) to iterate over code points. Python's
for ch in text already iterates over code points.
One symbol can contain several code points. What a reader sees as
one symbol is called a grapheme cluster. An accented letter can be stored as
one code point, such as é, or as e followed by a separate
accent code point. Therefore, café can contain four or five code points
even though it looks identical on screen.
Try café, emoji, Bengali text, and a combining accent in the
Four-way lab. Compare the character count with the UTF-8 byte
count to see the difference yourself.
Ready to move on when: you can explain why 👋 looks like
one symbol even though JavaScript reports "👋".length as 2.
Byte level
Computers store UTF-8 text as bytes. A byte is simply a number from
0 to 255, so there are exactly 256 possible byte values. A
byte-level tokenizer begins with one token for each of those values. This small base
vocabulary can represent any valid UTF-8 text, including unfamiliar names, emoji,
code, and mixed languages. If no larger token matches, the tokenizer can always fall
back to the original bytes instead of producing [UNK].
The disadvantage is that raw bytes can create long sequences. The letter
A uses one UTF-8 byte, but 👋 uses four. Without any learned
shortcuts, that one emoji would therefore require four byte tokens.
Traditional BPE vs byte-level BPE
Byte-level BPE does not replace the original BPE algorithm. Both versions begin with
basic symbols, count adjacent pairs, merge the most frequent pair, and repeat. The
difference is the starting unit: traditional character-level BPE learns from the
Unicode characters found in its training data, while byte-level BPE first converts
text to UTF-8 and starts from all 256 possible byte values.
That fixed byte vocabulary can represent every Unicode string, even when a particular
emoji, Hindi character, or Chinese character never appeared during training. Instead
of replacing unfamiliar text with [UNK] and losing information, the
tokenizer falls back to bytes and then uses the same BPE merges to combine frequent
byte sequences into larger tokens. This universal base vocabulary and efficient merge
process are why GPT-style tokenizers use byte-level BPE.
BPE starts with individual bytes and repeatedly joins byte sequences that appear
together often. Think of each learned merge as a shortcut. For example, the word
cat has these UTF-8 bytes, written in hexadecimal:
The final token still represents the same three bytes; BPE has only grouped them into
one frequently used piece. Decoding expands the piece back to
63 61 74, and UTF-8 turns those bytes back into cat.
The same idea applies to text that needs several bytes per visible symbol:
Whether the emoji becomes one token depends on the training corpus and vocabulary
budget. A common byte sequence is likely to receive a shortcut; a rare sequence may
remain split into several tokens. Either way, every byte is still representable.
GPT-style byte BPE. GPT-2 and OpenAI's tiktoken encodings use this
byte-based BPE idea. GPT-2 maps raw byte values to visible Unicode symbols while it
learns and stores merges, because many raw bytes are not printable. The visible
symbols are an internal representation; each learned token ultimately stands for one
byte or a sequence of bytes. Modern vocabulary sizes vary by model and include these
learned byte sequences plus special tokens.
Why not just use Unicode code points as the vocabulary?
This is the obvious question, and it deserves a real answer. Unicode already numbers
every character. Python even hands you the number for free with
ord("A") == 65. Why not declare “vocabulary = Unicode” and skip training a
tokenizer entirely?
Because it loses on both of the axes a tokenizer is judged on, at the same
time.
Problem 1 — far too many tokens
One code point per token means no compression whatsoever. The sentence
Tokenization is the first step of every language model. is 9 words and
55 code points, so it costs 55 tokens. A trained subword tokenizer
encodes the same sentence in about a dozen. Every token you spend is a position in the
context window and a step of attention — and attention cost grows quadratically with
sequence length, so a sequence roughly 4–5× longer is far more than 4–5× more
expensive.
Problem 2 — a huge and mostly dead vocabulary
A larger vocabulary requires the model to store and score more token IDs. This affects
both the beginning and the end of the transformer.
At the beginning: the input embedding table stores one learned
vector for every token ID. If the vocabulary has 50,000 tokens, it needs 50,000
rows. Each row has d_model numbers describing that token.
At the end: the output layer compares the transformer's current
hidden state with every possible next token. It produces one raw score, called a
logit, for each vocabulary entry.
Softmax: converts those logits into probabilities. For example,
cat: 0.60, dog: 0.25, and all other tokens sharing the
remaining probability.
The input table has shape vocab_size × d_model. The output projection has
the opposite shape, d_model × vocab_size, but both contain the same number
of values: vocab_size × d_model. Many models use weight tying,
which means these two layers share one set of parameters. Models without weight tying
pay this parameter cost twice.
Consider a small transformer with d_model = 768. One matrix needs:
number of parameters = vocabulary size × 768
50,257 × 768 = 38,597,376 parameters ≈ 38.6 million
The same calculation shows how quickly the cost grows:
Vocabulary
Size
Parameters in one matrix
What this means
Raw bytes
256
0.2 million
Very small vocabulary, but raw text needs many tokens
Example byte-level BPE
50,257
38.6 million
Larger vocabulary, but common byte sequences become shorter
Every assigned Unicode character
159,801
122.7 million
Large table with many characters a particular model may rarely use
The Unicode table in this example needs about three times as many
parameters as the BPE table. The model must also produce and compare 159,801 logits at
every output position instead of 50,257. A bigger vocabulary can shorten token
sequences, but it also makes the embedding and output layers wider. Tokenizer design
is a balance between those two costs.
Worse, most of that table is dead weight. An English corpus touches maybe a hundred
distinct code points; the remaining ~159,700 rows receive almost no gradient and stay
essentially random. You would be paying for capacity you never train.
Problem 3 — it is still not closed
You might accept the cost for the guarantee that nothing is ever unknown. You do not
get that guarantee. A character-level vocabulary is built from the characters in
your training corpus, so any character outside it is still [UNK]. And
Unicode itself keeps growing — version 17.0 added 4,803 characters and 4 new scripts.
A vocabulary pinned to a Unicode version goes stale.
The verdict. Unicode-character tokenization gives you long sequences
and a large vocabulary and still no coverage guarantee — the worst
corner of the tradeoff. Bytes fix the second and third problems outright, and BPE on
top of bytes fixes the first.
Character level vs byte level, side by side
Character level (Unicode)
Byte level (UTF-8)
Atomic unit
One Unicode code point
One byte, 0–255
Base vocabulary size
Open-ended — up to ~160,000 and growing every release
Fixed at exactly 256, forever
Unseen input
[UNK] — the spelling is lost
Impossible; every byte is already in the vocabulary
Reuse across scripts
None. é, e, and আ are unrelated IDs with no shared structure
Heavy. The same 256 atoms are reused by every script on earth
Cost of one character
Always 1 token
1–4 tokens before merges; usually back under 1 after BPE
Rows that stay untrained
The vast majority, on any single-language corpus
Essentially none — all 256 bytes get used
The byte column’s “1–4 tokens per character” looks like a loss, and on raw bytes it is:
বাংলা ভাষা is 10 code points but 28 UTF-8 bytes. But those bytes are only
the starting alphabet. BPE then merges the recurring byte sequences of Bengali
into single pieces, and the sequence collapses back down. You get the coverage of bytes
with the length of subwords.
Byte fallback: the hybrid most models actually ship
You do not have to choose one or the other. Byte fallback keeps a
normal subword vocabulary for everything common, and drops to raw bytes only for
things that vocabulary cannot express:
Text that matches a learned piece → encoded as that piece, as usual.
Anything else — a rare script, a brand-new emoji, a corrupted byte — is emitted as
individual byte tokens such as <0xF0><0x9F><0x98><0x8A>.
The result is that [UNK] never appears and decoding is always exact, at
the price of a few extra tokens in rare cases. This is SentencePiece’s
byte_fallback option, and it is what the Llama tokenizers use. It is the
practical answer to “what if something cannot be represented?” — the fallback is the
raw byte.
UTF-8 inspector
Bytes
Grapheme
Code point(s)
CP #
UTF-8 bytes
Byte #
Leave when: you can read one row of the inspector and separate
grapheme / code point / byte counts.
Vocabulary size, context length, and sequence length
Three numbers that beginners routinely merge into one. They are independent, they are
set at different times, and confusing them produces confident nonsense about what a
model can do.
Does “1B” mean one billion vocabulary tokens?
No. In a model name such as Llama 3.2 1B, the 1B means the
neural network has roughly one billion learned parameters. Likewise,
7B means about seven billion parameters and 70B means about
seventy billion. The letter B means billion; it does not describe
the tokenizer vocabulary.
A parameter is a number adjusted while the model is trained. Together,
billions of these numbers let the network learn language patterns and store useful
information. Parameters are spread across the embedding layer, attention layers,
feed-forward layers, and other parts of the model.
Number
What it measures
Example meaning
1B, 7B, 70B
Total model parameters
Approximately 1, 7, or 70 billion learned numbers
vocab_size: 128256
Tokenizer vocabulary
128,256 different token IDs are available
context_length: 131072
Context capacity
Up to 131,072 token positions can fit in one request
These numbers do not have to grow together. A two-billion-parameter model can use a
vocabulary of 32,000 tokens, 128,000 tokens, or another size chosen by its designers.
To find the real vocabulary size, check the model configuration or tokenizer files
for a field such as vocab_size.
How vocabulary contributes to parameters. If a model has a vocabulary
of V tokens and each token embedding contains D numbers, its
embedding table contains V × D parameters. A larger vocabulary therefore
increases the model's parameter count, but it does not determine the model's name.
Most parameters are usually in the transformer blocks, although embeddings can take a
noticeable share of a small model with a large vocabulary.
model name: 7B → total learned parameters
tokenizer config: vocab_size → number of available token IDs
model config: context_length → maximum token positions per request
What a vocabulary size of 50,257 actually means
GPT-2’s vocabulary size is 50,257. That number is not arbitrary — it decomposes
exactly:
Component
Count
What it is
Base byte tokens
256
Every possible UTF-8 byte — the atoms nothing can fall through
Learned merges
50,000
The merge budget BPE was given during training
Special token
1
<|endoftext|>
Total
50,257
Read this carefully. A vocabulary size of 50,257 means there are
50,257 unique token pieces — not 50,257 words. Most entries
are not words at all: single bytes, punctuation, whitespace runs, and word fragments
like ing, tion, or the (with its leading space
attached).
And because pieces combine, the set of strings the tokenizer can
represent is not capped at 50,257. A word that is not in the vocabulary is simply split
into smaller pieces that are, then reassembled on decode:
unhappiness → un + happiness 2 pieces
tokenization → token + ization 2 pieces
Kubernetes → Kub + ernet + es 3 pieces
zyxwvu → z + y + x + w + v + u 6 pieces (worst case: one per byte)
Sequences of pieces multiply. With 50,257 pieces, even two-piece combinations already
exceed two billion distinct strings, and there is no limit on sequence length.
Therefore BPE can represent vastly more than 50,257 words — in fact
every string that can be written in UTF-8, including words that will be invented next
year. The vocabulary size caps how efficiently text is encoded, never
whether it can be encoded.
That is the whole reason bigger vocabularies keep shipping: more pieces means fewer
pieces per sentence, not more words covered. Coverage was already total.
Tokenizer
Algorithm
Vocabulary size
BERT base uncased
WordPiece
30,522
Llama 2
SentencePiece BPE
32,000
GPT-2 (r50k_base)
Byte-level BPE
50,257
GPT-4 (cl100k_base)
Byte-level BPE
100,277
Llama 3
Byte-level BPE
128,256
GPT-4o (o200k_base)
Byte-level BPE
200,019
Vocabulary size is not context length
These answer two completely different questions:
Vocab size
How many different tokens exist. The size of the dictionary. Fixed when
the tokenizer is trained, and frozen for the life of the model — changing it means
retraining the embedding table.
Context length
How many tokens the model can process at once. The number of positions in
the window. Set by the model architecture and training, and has nothing to do with
how many distinct tokens exist.
Sequence length
How many tokens your particular input came out as. Varies per input, and
must fit inside the context length.
GPT-2 makes the independence obvious: vocabulary 50,257, context
length 1,024. The dictionary has fifty thousand entries; the model can
read a thousand tokens at a time. Neither number constrains the other.
Analogy. Vocabulary size is how many words are in your dictionary.
Context length is how many words you can hold in your head at once. Learning more
words does not enlarge your working memory — but it does let you say the same thing in
fewer words, so more meaning fits.
That last clause is the one real connection between them. A larger vocabulary does not
change the context length, but it lowers the token count of the same text, so
more text fits in the same window. This is why o200k_base encodes
non-English text in noticeably fewer tokens than cl100k_base despite the
window being a property of the model, not the tokenizer.
See the tradeoff
Same sentence, four granularities. The bars are token counts. This is the tradeoff
every tokenizer paper is negotiating.
Length comparison
Counts
Why don’t modern models use word-level tokenization if it uses fewer tokens?
Word-level tokenization looks efficient in this chart because every word in the sample
sentence is familiar. In real text, that advantage quickly breaks down.
An unknown word loses its identity. A word vocabulary is frozen
before model training. If COVID-19, useState, a username, a
rare name, or a misspelling is missing, the tokenizer returns [UNK]. The
model does not receive the word's spelling, so it cannot see which known parts it
contains. A subword tokenizer can preserve that information using smaller pieces;
byte-based subwords can always fall back to bytes.
Related word forms become unrelated IDs. A word tokenizer stores
play, plays, played, and playing as
four independent entries. Subwords can reuse pieces such as play,
ed, and ing. This matters even more for languages with many
word forms and productive compounds, including Turkish, Finnish, Arabic, and
German.
Better coverage requires an enormous vocabulary. Adding more whole
words reduces [UNK], but every entry needs an embedding and an output
score. A vocabulary with hundreds of thousands of words makes those layers larger
and makes next-token scoring more expensive. Rare entries may still appear too few
times to learn useful representations.
Multilingual text does not fit neatly into one word list. A shared
word vocabulary would need entries for words from every supported language, plus
mixed-language text, emoji, and code. Subword and byte-based tokenizers reuse a
smaller set of pieces across all of them.
The token saving is smaller than it first appears. Keeping a common
word as one token instead of two saves one position. That modest saving is not worth
losing every rare or newly invented word. Subwords accept slightly longer sequences
for common text in exchange for reliable coverage of the long tail.
training vocabulary contains: play, football, the
word level
playing → [UNK]
footballer → [UNK]
useState → [UNK]
subword level
playing → play + ing
footballer → football + er
useState → use + State
The real tradeoff. Word-level tokenization gives the shortest sequence
for known words, but fails on unknown words and needs a vocabulary that grows with
every language and word form. Subwords use a few more token positions, but keep the
vocabulary manageable, reuse structure, and preserve unfamiliar text. That is why
modern language models generally choose subword or byte-based subword tokenization.
Leave when: you can say which bar is not a production
tokenizer, and explain to someone else why a 50,257-piece vocabulary can represent far
more than 50,257 words while saying nothing at all about context length.
Five printable schematics (Token Lab cyan on light paper). Open full-page for
projectors — better than mermaid fences when Jupyter does not render diagrams.
The universal standard that assigns a unique code point to every
character: A → U+0041, 😊 →
U+1F60A. It defines which characters exist, not how they are
stored.
Code point
One numbered entry in the Unicode catalogue, written U+XXXX in hex.
159,801 of the 1,114,112 possible code points are assigned as of Unicode 17.0.
UTF-8
The encoding that turns Unicode code points into bytes for storage and
transmission. ASCII characters take 1 byte; others take 2–4. Unicode defines the
characters, UTF-8 defines the bytes.
Byte
A value 0–255. There are exactly 256 of them and that number never grows — the appeal of byte-level tokenizers.
Grapheme
What a reader calls “one character” on screen. May be several code points (a letter plus a combining accent, or an emoji sequence).
Byte fallback
Emitting raw byte tokens such as <0xF0> for anything the subword
vocabulary cannot express. Removes [UNK] entirely. Used by the Llama
tokenizers via SentencePiece byte_fallback.
Vocab size
How many distinct token pieces exist. GPT-2: 50,257 = 256 bytes + 50,000 merges + 1
special token. Pieces, not words — they combine, so coverage is unlimited.
Context length
How many tokens the model can process at once. A property of the model, entirely
separate from vocab size.
Logits
The raw scores the final layer produces — one per vocabulary entry, before softmax.
A bigger vocabulary means a wider logit vector at every position.
Embedding
The vector a token ID is looked up as. The lookup table is vocab_size × d_model, which is why vocab size costs parameters.
Greedy
Takes the best option available at the current step, with no lookahead and no backtracking. BPE merge selection is greedy.
OOV / [UNK]
Out of vocabulary. English: unknown word. বাংলা: অজানা শব্দ — the model sees a hole.
Type
A distinct vocabulary entry (the string), not each occurrence in a sentence.
Merge
BPE / WordPiece training step that glues two adjacent symbols into one new piece.
##
WordPiece continuation mark. ##est continues a word; it does not start one.
</w>
BPE end-of-word mark. Distinguishes the st in star from the end of widest.
▁ (U+2581)
SentencePiece space mark — not a plain underscore. Decoding turns each ▁ into a space.
NFKC
Unicode normalization form. Makes compatibility-equivalent characters look the same before training.
Next: four tokenization playbooks
The labs above are intuition. The notebooks build the machines from scratch on the
Sennrich toy (or a tiny Viterbi demo), then repeat the same idea on
data/tiny_corpus.txt with the current PyPI libraries via uv.
The fourth playbook compares traditional and byte-level BPE directly.
01_bpe.ipynb — Byte Pair EncodingFrequent pair merges, HuggingFace BpeTrainer, then tiktoken as production BPE.Open BPE Explained playbook in Colab
02_wordpiece.ipynb — WordPieceLikelihood score instead of raw frequency, ## continuation, BERT-style trainer.
04 — Byte-level BPE vs traditional BPESame merge algorithm, different base symbols: observed Unicode characters vs a universal 256-byte vocabulary.Open byte-level BPE comparison playbook in Colab
Tokenize the same real prompt with three model tokenizers and compare the results:
BERT: use its WordPiece tokenizer and look for continuation pieces
marked with ##.
GPT-4o: use its OpenAI tiktoken encoding and inspect how
spaces, punctuation, and UTF-8 text are grouped.
T5: use its SentencePiece tokenizer and look for the
▁ symbol that represents a word-starting space.
For each tokenizer, record the token pieces and token count. Then explain why the
counts differ even though the input text is identical. Use each model's real tokenizer:
tiktoken is appropriate for GPT-4o, but BERT and T5 require their own
tokenizer implementations.