Posted in

What Is Tokenization in AI? How Do LLMs Turn Text Into Tokens

AI tokenization process showing text being converted into tokens for an LLM
A visual representation of how text is converted into tokens before being processed by a language model.

When you ask a question in an AI chatbot, the model does not read your sentence as you would read it. Before the language model can process your request, your sentence is split into smaller meaning units called tokens. This operation is called tokenization. It is one of the first steps in the way modern language models process human languages.

Tokenization is a part of the natural language processing workflow, needed to make text understandable to a computer program. A token can be a whole word, a part of it, or even punctuation or other symbols. It all depends on the model and its tokenizer.

What Is a Token?

A Token is a chunk of text that the language model processes. Typically, a token is not necessarily a word since some words may get split into smaller parts. As expected, words can also contain multiple tokens. Consider the sentence “AI is changing technology.” A tokenizer could break the sentence into multiple tokens. Different models should tokenize the same text in various ways. Some language models will use a different set of tokens to represent the same word.

Most modern language models use a smart method called a sub word tokenization. The idea is that while common words should be kept whole, less common words should be split into multiple parts. So, a word like “unhappiness” may be separated into multiple smaller units with another word serving as a token bridge. The language model does not need to store every existing word as a separate token as you would have to type “unhappiness” for it to recognize the word. Instead, it can combine smaller pieces to build up the word.

How Does Tokenization Work?

The process can be demonstrated by walking through the steps taken by an AI after a sentence is entered into the system. Breaking the text down into smaller elements called tokens makes up the first step. Then, these tokens are converted to their respective token IDs-which essentially serve as numbers corresponding to entries on a model’s vocabulary. The model can later interpret these converted tokens, called embeddings, and proceed to execute the process of passing them to its Transformer layer. Hence, the overall process can be summarized as:

Text → Tokens → Token IDs → Embeddings → Transformer → Output

This is where tokenization connects with large language models and Transformer models. Tokenization prepares the text for the numerical processing that happens inside the model. It does not give the model an understanding of the sentence by itself. Instead it creates the pieces that the later stages of the model can work with.

Why Do Models Split Words Into Smaller Pieces?

Language has an extraordinary amount of words. Besides that, there are names; technical terms and neologisms; and spelling variations and subtleties. A model which exclusively used complete words would require a vast amount of words to form its vocabulary. Subword tokenization offers a solution in that common words can be represented as themselves while uncommon words are split into smaller pieces that already exist. Methods such as Byte Pair Encoding, WordPiece, and Unigram differ in how they generate these smaller units.

This is also the reason why two tokens can represent the same sentence differently. One token keeps a word whole while another splits it into multiple parts; there is no absolute rule which states that one word is always equal to one token. Tokens Are Not the Same as Words, It is important that the distinction is made, especially when using an LLM. While people generally measure text in words, a language model will use tokens. A document which contains 1,000 words does not necessarily contain 1,000 tokens as well. Some words have to be represented with more than one token. Additionally, punctuation can also take up tokens.

The number of tokens depends on both the model and the tokenizer. It is important to remember this because many language models base the limits of their context window on tokens. A long document will consume the context of a model faster than one would expect. The same applies to the output of the language model. An LLM will generate its output as a series of tokens, which are subsequently converted to words. Everything that a person sees as one coherent sentence has to be generated step by step by the model.

Tokenization and Embeddings

Tokenization and embeddings are tightly related but different processes. First, tokenization splits text into tokens and encodes them with particular ID’s. Then, embeddings convert these tokens into vectors by encoding them as numbers. These representations enable the Transformer to perform self-attention. Self-attention operates not on text tokens but on representations learned by encoder layers. In other words, the encoder converts input text into numerical form to help the Transformer network identify patterns and relationships between tokens.

Thus, it is essential to understand that tokenization is just a part of the whole process. Text inputs are converted into tokens using a tokenizer. Then, tokens are turned into numerical representations. Finally, encoded tokens are used to perform computations that help the network comprehends the input.

Why Tokenization Matters

While tokenization may seem like a small technical detail, it plays an essential role in the functioning of a language AI. It is the process, which makes a practical implementation of the model possible by providing a sensible representation of the consumed text in a form, which can be used by an algorithm. After the input text is tokenized, the tokens can be converted to embeddings, which are then processed by a transformer. The process can be generally simplified with the following diagram:

Text → Tokens → Token IDs → Embeddings → Transformer → Output

It becomes apparent that the process is not complex at all – that is why grasping the details of tokenization is crucial to understanding how the modern language models work. Having learned what tokens are, you can better understand the inner workings of any AI system, which utilizes them. Next time you hear about a token limit in a large language model, you will be able to explain the concept to someone else. A sentence, which you input in the chat is first converted to tokens. The tokens are then transformed into numerical representations – embeddings. The language model, which consists of multiple transformers, processes the embeddings and outputs the result, which you see in your browser.

FAQ’s

What is tokenization in AI?

It is the process of segmenting text / input into smaller pieces tokens which can be further processed by the AI model.

Is one token one word?

Not necessarily. It can be a part of a word, a whole word, or a punctuation mark – anything can be tokenized

Why do LLMs use tokens?

The models process numeric data, so the tokens serve as the basis for the numeric representation of text.

Why is it important how many tokens are there?

Simply because more tokens mean more text. Also, the context length is often counted in tokens.

Do different models use the same tokenizer?

No, they may have different tokenizers and vocabularies. So the same text can have a different number of tokens depending on the model.

What happens to tokens after they are created?

They are mapped to their numeric IDs and embedded – their vector representations are created, which the model further processes through its layers.

Sources

Krish Shrestha, founder and editor of LegacyVia

Written & Researched by

Krish Shrestha

Krish Shrestha is the founder and editor of LegacyVia. He researches and writes about AI and technology with a focus on understanding how new technologies work and explaining them in a clear, practical way.

Leave a Reply

Your email address will not be published. Required fields are marked *

Krish Shrestha
Founder & Editor Krish Shrestha Founder and editor of LegacyVia, an independent publication covering AI and technology. He researches, writes, and maintains every article on the site.
TRENDING