Flowchart · Token Lab / 01

Tokenizer pipeline

Six jobs in order. Skip one and lengths, unknowns, or multilingual text break later.

01 Raw text 02 Normalize 03 Pre-tokenize 04 Model 05 IDs 06 Decode Unicode NFKC / case split rules word·sub·char·byte integers string back LEGEND Focal stage (granularity choice) Start / end Step ← All diagrams