Skip to content
Fundamentals Glossary

Tokenization

The process of splitting text into tokens before it is fed to a language model.

Last updated

Tokenization converts raw text into the units a model can process. Most systems use subword algorithms such as byte-pair encoding, which balance vocabulary size against the ability to represent rare words. Good tokenization keeps common sequences intact while still handling unusual spellings and code. It is the first step in every NLP pipeline.

Related terms

Read more about Tokenization

Keep exploring

Browse the full AI glossary or compare AI tools that use this technology.

Report an issue with this page