Back to BlogNLP

Tokenization in NLP

1. Introduction

image.png

The Nature of Tokenization

In the world of Natural Language Processing (NLP), one of the largest barriers between humans and computers is the way information is received. Humans read and understand text as a continuous stream of characters, words, and context, where meaning is inferred through experience and abstract thinking. By contrast, machine learning models and artificial intelligence are completely blind to raw text. Their core architecture, from traditional neural networks to state-of-the-art Transformer architectures, can only receive and compute on numerical matrices and mathematical vectors. To close this perceptual gap, the first, most fundamental, and most decisive preprocessing step is Tokenization.

Put in the most intuitive and accessible terms, Tokenization is the process of dissecting a piece of raw text and splitting it into smaller, manageable fragments called "tokens." These fragments are not limited to a single format; they can be complete words, parts of words (prefixes, suffixes, roots — also known as subwords), or even the individual characters that make up a language. The ultimate goal of this process is to convert a complex, polysemous linguistic message into an array of discrete elements. After the split, each unique token is looked up in the system's "vocabulary" and assigned a unique identifier (ID). Only from those IDs can the computer map them into vector space (embedding space) and begin learning semantic relationships.

In academic and in-depth technical settings, drawing a sharp distinction among foundational concepts is essential for designing systems accurately. Learners need to thoroughly understand the difference among three terms that are often confused: Token, Type, and Term.

  • A "Token" is a concrete physical entity that appears in a piece of text.
  • A "Type" represents the class containing all tokens that share exactly the same character sequence.
  • A "Term" is a type that has undergone normalization (for example, converting everything to lowercase and removing morphological variants) and is stored formally in the information system's dictionary.

To illustrate, if we consider the classic quotation "to sleep perchance to dream," this text contains a total of 5 distinct tokens. However, because the token "to" appears twice, we have only 4 types. If the system decides to drop "to" because it is a stop word that does not carry much classificatory meaning, the final dictionary will retain only 3 core terms: "sleep", "perchance", and "dream".

The Importance of Tokenization in the NLP Pipeline

The success or failure of every NLP pipeline, from basic sentiment classification tasks to pretraining Large Language Models (LLMs) with billions of parameters, is determined to a very large degree by the quality of the Tokenization strategy. No AI algorithm, however excellent, can compensate for text input that has been split in a way that distorts meaning. This importance shows up clearly in the three structural aspects below.

First, Tokenization directly shapes the size and geometry of the input data space by deciding Vocabulary Size. A naive word-splitting method can produce a huge vocabulary containing millions of rare or misspelled tokens. This leads to a dimensional explosion of the embedding matrix. A typical example from the past: when Transformer XL used whitespace- and punctuation-based word splitting, it had to carry a vocabulary of up to 267,000 words. As a result, the model possessed extremely large embedding matrices at both the input and output layers, exponentially increasing both time complexity and computational memory. Thanks to the evolution of Tokenization, modern transformer models today rarely need a vocabulary larger than 50,000 tokens, yet still ensure comprehensive language coverage.

Second, Tokenization is the key to solving the Out-of-Vocabulary (OOV) problem and improving the Generalization of machine learning models. In the real world, language constantly changes with new slang, typing errors, or complex morphological variants. If a system only recognizes whole words, it will be completely paralyzed when it encounters a word that never appeared in the training data. By splitting vocabulary into morphologically meaningful subwords, Tokenization lets the model recombine basic units to understand brand-new words, thereby maintaining system stability in the noisy data environment of the real world.

Third, Tokenization directly affects information-transfer efficiency and input Sequence Length. Transformer architecture has quadratic computational complexity O(N2)O(N^2) with respect to token sequence length NN. If Tokenization cuts the text too finely (for example, into individual letters), the input sequence becomes extremely long, draining GPU memory and reducing the model's ability to capture long-range dependencies in a sentence. Conversely, if the cut is too coarse, semantic information is packed too tightly and becomes hard to analyze. Finding the ideal balance is the core art of this preprocessing step.

Visualizing the Data Transformation Process

To help learners picture the metamorphosis of data as thoroughly as possible, let us look at a practical example illustrating the transformation from raw text into a format a computer can absorb.

Consider the following input text: "Don't you love Transformers? We sure do.".

If we apply the most primitive whitespace-splitting method, the result is an array:

["Don't", "you", "love", "🤗", "Transformers?", "We", "sure", "do."]

At a glance, this result looks reasonable. However, from a semantic-analysis perspective, it reveals fatal gaps. Tokens such as "Transformers?" and "do." are stuck to the question mark and the period. That means the model will treat "Transformer" (without punctuation) and "Transformers?" (with punctuation) as two completely different entities, wasting vocabulary memory and polluting vector space.

A refined, modern Tokenization pipeline will process the sentence far more sharply. It recognizes grammatical rules, separates punctuation, preserves emoji, and handles contractions.

A complete Tokenization sequence will resolve the text into:

["Don", "'", "t", "you", "love", "🤗", "Transformers", "?", "We", "sure", "do", "."]

Finally, the system looks this array up in its internal vocabulary and converts it into a sequence of identifiers (Input IDs). Through the lens of this process, an English sentence with complex emotional coloring has been condensed into a compact mathematical tensor, ready to be fed into a deep neural network. The emoji is not discarded as garbage; it is treated as a token with independent sentimental value, contributing to the model's sentiment-analysis capability.

image.png

2. Core Component

The development of Computer Science in general and NLP in particular has witnessed the birth, competition, and replacement of many Tokenization strategies. Each strategy arose to resolve the deadlocks of the previous generation, especially the problem of balancing language diversity against hardware resource limits. To master the system, learners need a deep understanding of how each mechanism works, as well as a multidimensional analysis of the strengths and weaknesses of the three core Tokenization schools below.

image.png

2.1. Word-level Tokenization (splitting on natural boundaries)

True to its name, Word-level Tokenization approaches text in the way most natural to human thinking: using physical boundaries already present in the text, mainly whitespace and punctuation, as cut points. This method dominated the early era of NLP and still retains value in small-scale applications today.

The mechanism is relatively intuitive. When it receives an input string, the algorithm scans and cuts the string at every position that contains a space, a tab, or a newline. More sophisticated versions (such as Rule-based tokenization or Regular Expressions) are programmed with additional rules to separate commas and periods and to handle common English contractions.

However, that simplicity brings serious limitations when deployed at large scale or with diverse data. The greatest advantage of Word-level is its ability to preserve the full meaning of each word without breaking morphological structure, together with extremely low computational cost and easy deployment at the infrastructure level. Yet its drawbacks are destructive for deep models.

First, this method directly leads to Vocabulary Explosion. In any language, the number of words is unbounded because of prefixes, suffixes, and human linguistic creativity. If every variant of a word (for example: "run", "runs", "running", "ran") is treated as an independent token, vocabulary size will swell out of control, far beyond hundreds of thousands. This places enormous pressure on the memory of the weight matrix.

image.png

The second direct consequence is the Out-of-Vocabulary (OOV) problem. Faced with data containing typos, social-media slang, or a newly appearing technical term, a Word-level Tokenizer is completely helpless and must assign them an <UNK> (Unknown) label. Dense occurrences of the <UNK> token blind the model and strip away its ability to infer context. This method fails completely and entirely destroys the semantics of the text.

image.png

2.2. Character-level Tokenization (basic character splitting)

To thoroughly solve the nightmare named OOV and the explosion of vocabulary memory, researchers went to the opposite extreme: Character-level Tokenization. This method decomposes text down to the most microscopic level, treating each individual character (alphabetic letters, digits, symbols, punctuation) as an independent token.

image.png

This mechanism offers an unmatched advantage: vocabulary size is reduced to a minimum, typically only a few hundred tokens for most languages (depending on the supported Unicode encoding). More importantly, the OOV problem is eliminated entirely. Any strange word or spelling error is still composed of basic characters, so the computer can always represent them as a valid sequence of IDs.

image.png

Yet the price of this memory optimization is a sacrifice in semantic power and computational performance. The largest drawback is Loss of Word Meaning. A letter "a" or "b" standing alone carries no lexical meaning. The responsibility of connecting these meaningless characters into a meaningful structure is pushed entirely onto the hidden layers of the machine learning model. The model must be deeper, need more data, and take longer to converge in order to learn on its own that the character cluster "a-p-p-l-e" put together forms the word apple.

In addition, Character-level produces an input sequence of enormous length. A sentence of only 20 words can become a sequence of more than 100 character tokens. As mentioned in the overview, for models based on Transformer architecture, computational cost grows exponentially with sequence length O(N2)O(N^2). Feeding an overly long sequence into the model not only slows inference but also quickly exhausts GPU VRAM, making this method unusable for tasks that require processing long text.

image.png

2.3. Subword-level Tokenization

Recognizing the inherent limits of the two methods above, AI architects created Subword-level Tokenization — a design philosophy of perfect compromise that serves as the backbone of the rise of LLMs such as GPT, BERT, and T5. The philosophy of the method is to preserve frequent words intact, and to decompose rare words into subword units that carry morphological meaning.

image.png

With this strategy, vocabulary size is tightly controlled at an optimal level (typically ranging from 30,000 to 64,000 tokens), small enough to compute efficiently but also large enough not to turn text into overly long sequences. Words that have never appeared (OOV) are not turned into a meaningless <UNK>; instead, the model interprets them by assembling prefixes, roots, and suffixes.

To help learners grasp this deeply, below is a detailed analysis of the three most modern and powerful Subword algorithms in use today.

A. Byte Pair Encoding (BPE)

Byte-Pair Encoding (BPE) was introduced in the paper Neural Machine Translation of Rare Words with Subword Units (Sennrich et al., 2015). BPE relies on a pre-tokenizer to split the training data into words. The pre-tokenization process can be as simple as splitting on whitespace (space tokenization), as in GPT-2 and RoBERTa. More advanced pre-tokenization methods include rule-based tokenization, for example XLM and FlauBERT (using the Moses toolkit for most languages), or GPT (using spaCy and ftfy), in order to count the frequency of each word in the training corpus.

After pre-tokenization, a set of unique words is created and the frequency of each word in the training data is also determined. Next, BPE creates a base vocabulary consisting of all symbols that appear in that set of unique words, then learns merge rules to form a new symbol from two symbols in the base vocabulary. This process is repeated until the vocabulary reaches the desired vocabulary size. Note that the desired vocabulary size is a hyperparameter that must be determined before tokenizer training begins.

As an example, suppose that after pre-tokenization, the following set of words together with their frequencies has been determined:

("hug", 10), ("pug", 5), ("pun", 12), ("bun", 4), ("hugs", 5)

Therefore, the base vocabulary is ["b", "g", "h", "n", "p", "s", "u"]. By splitting all words into symbols from the base vocabulary, we have:

("h" "u" "g", 10), ("p" "u" "g", 5), ("p" "u" "n", 12), ("b" "u" "n", 4), ("h" "u" "g" "s", 5)

The BPE algorithm then counts the frequency of every possible symbol pair and selects the pair that occurs most often. In the example above, "h" followed by "u" occurs 10 + 5 = 15 times (10 times in the 10 occurrences of "hug", 5 times in the 5 occurrences of "hugs"). However, the most frequent symbol pair is "u" followed by "g", for a total of 10 + 5 + 5 = 20 times. Therefore, the first merge rule the tokenizer learns is to group together every "u" immediately followed by "g". Next, "ug" is added to the vocabulary. The set of words then becomes:

("h" "ug", 10), ("p" "ug", 5), ("p" "u" "n", 12), ("b" "u" "n", 4), ("h" "ug" "s", 5)

BPE continues to identify the next most common symbol pair. That is "u" followed by "n", occurring 16 times. "u" and "n" are merged into "un" and added to the vocabulary. The next most frequent symbol pair is "h" followed by "ug", occurring 15 times. Once again, this pair is merged and "hug" is added to the vocabulary.

At this stage, the vocabulary is ["b", "g", "h", "n", "p", "s", "u", "ug", "un", "hug"] and our set of unique words is represented as:

("hug", 10), ("p" "ug", 5), ("p" "un", 12), ("b" "un", 4), ("hug" "s", 5)

Suppose Byte-Pair Encoding (BPE) training stops at this point; the learned merge rules are then applied to new words (as long as those new words do not contain symbols that are not already in the base vocabulary). For example, the word "bug" would be tokenized as ["b", "ug"], but "mug" would be tokenized as ["<unk>", "ug"] because the symbol "m" is not in the base vocabulary. In general, individual letters such as "m" are usually not replaced by the "<unk>" (unknown) symbol because training data typically includes at least one occurrence of each letter, but this happens very easily with special characters such as emojis.

As mentioned earlier, vocabulary size — that is, base vocabulary size + the number of merges — is a hyperparameter that must be chosen. For example, GPT has a vocabulary size of 40,478 because they used 478 base characters and decided to stop training after 40,000 merges.

B. WordPiece

Developed by Google and made famous worldwide by the success of BERT, WordPiece shares the same philosophical foundation as BPE but carries a more optimal mathematical refinement.

image.png

The core difference lies in the selection criterion for merging. While BPE relies purely on raw frequency counts, WordPiece applies probability theory. It computes and selects the character pair that maximizes the Likelihood of the entire training data when that pair is merged. Specifically, it evaluates the probability of the combined pair compared with the probability of the two characters standing completely independently. This helps WordPiece make splits that are more phonetic and morphological.

The identifying feature of WordPiece is the use of a special prefix symbol (usually ##) to mark subwords as auxiliary pieces that must be attached to the word immediately before them. A complex proper noun or a newly coined word can be skillfully represented by WordPiece as ["pro", "##per", "##noun"]. In practical deployments, WordPiece has been shown to cover morphological variation, creative spelling, and multilingual context excellently, significantly reducing the number of OOV tokens compared with older methods.

C. Unigram Language Model Tokenization

Unigram is a subword tokenization algorithm introduced in the paper Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates (Kudo, 2018). In contrast to BPE or WordPiece, Unigram initializes its base vocabulary with a large number of symbols and gradually prunes symbols to obtain a smaller vocabulary. For example, the base vocabulary may correspond to all pre-tokenized words and the most common substrings. Unigram is not used directly for any model in the transformers library, but it is used in combination with SentencePiece.

At each training step, the Unigram algorithm determines a loss function (usually defined as log-likelihood) on the training data, based on the current vocabulary and a unigram language model. Then, for each symbol in the vocabulary, the algorithm computes how much the total loss would increase if that symbol were removed from the vocabulary. Unigram then removes pp percent (with pp typically 10% or 20%) of the symbols with the smallest loss increase, that is, the symbols that affect total loss on the training data the least. This process is repeated until the vocabulary reaches the desired size. The Unigram algorithm always keeps the base characters so that every word can still be tokenized.

Because Unigram does not rely on merge rules (unlike BPE and WordPiece), the algorithm has several different ways to tokenize new text after training. For example, if a trained Unigram tokenizer has a vocabulary such as:

["b", "g", "h", "n", "p", "s", "u", "ug", "un", "hug"]

The word "hugs" can be tokenized as ["hug", "s"], ["h", "ug", "s"], or ["h", "u", "g", "s"]. So which one should we choose? Unigram stores the probability of each token in the training data alongside the vocabulary, so the probability of every possible tokenization can be computed after training. In practice, the algorithm typically simply chooses the most likely tokenization, but it also provides the ability to sample a feasible tokenization based on those probabilities.

These probabilities are defined by the loss function the tokenizer was trained with. Suppose the training data consists of words x1,…,xNx_1, \dots, x_N and the set of all possible tokenizations of a word xix_i is defined as S(xi)S(x_i); then the total loss LL is defined as:

L=−∑i=1Nlog⁡(∑x∈S(xi)p(x))L = -\sum_{i=1}^{N} \log \left( \sum_{x \in S(x_i)} p(x) \right)
Analytical CharacteristicWord-levelCharacter-levelSubword-level (BPE/WordPiece)
Algorithmic ComplexityVery LowVery LowHigh (requires a vocabulary-training stage)
OOV Risk LevelVery High (blinds the model)NonexistentMedium to Low
Semantic PreservationPerfect (each word is a concept)Very Poor (meaningless characters)Good (keeps the root, splits suffixes)
Sequence LengthShortExtremely LongMedium (optimal balance)
Recommended ApplicationsBasic analysis, narrow English dataDNA analysis, gene sequences, spelling errorsMandatory standard for every LLM, machine translation

3. Implementation

image.png

3.1. Basic Tokenization (using NLTK and SpaCy for English)

In the context of small- to medium-scale English data analysis, NLTK (Natural Language Toolkit) and spaCy are two monuments that cannot be ignored. They operate on sophisticated linguistic rule sets, whitespace splitting, and excellent exception handling.

# Environment setup instructions:
# pip install nltk spacy
# Need to download the language model for spaCy: python -m spacy download en_core_web_sm

import nltk
from nltk.tokenize import word_tokenize
import spacy

# Download NLTK's basic sentence-splitting resource (only needs to run once)
nltk.download('punkt', quiet=True)

def basic_english_tokenization(text: str):
    """
    Demonstrate basic Tokenization with NLTK and SpaCy.
    Purpose: Observe how the algorithm handles contractions, punctuation, and special symbols.
    """
    print(f"--- INPUT TEXT (RAW TEXT) ---")
    print(f"{text}\n")
    
    # 1. Implement NLTK Tokenization
    # NLTK uses a basic machine learning model (Punkt) to recognize word and punctuation boundaries
    nltk_tokens = word_tokenize(text)
    print(f"--- NLTK TOKENIZATION RESULT ---")
    print(f"Total Tokens: {len(nltk_tokens)}")
    print(f"Token Array: {nltk_tokens}\n")
    
    # 2. Implement SpaCy Tokenization
    # SpaCy performs comprehensive syntactic analysis, loading a compact English language model
    nlp = spacy.load("en_core_web_sm")
    doc = nlp(text)
    
    # Extract the text form of each token from the Doc object
    spacy_tokens = [token.text for token in doc]
    
    print(f"--- SPACY TOKENIZATION RESULT ---")
    print(f"Total Tokens: {len(spacy_tokens)}")
    print(f"Token Array: {spacy_tokens}\n")

# Test scenario containing grammatical traps: contractions (doesn't, O'Neill), symbol ($)
sample_text = "Mr. O'Neill doesn't think the new AI model is worth $100.50!"
basic_english_tokenization(sample_text)

When running the code above, learners will notice that both NLTK and spaCy demonstrate intelligence far beyond the ordinary split() function.

The word "doesn't" is correctly split into the root "does" and the negative suffix "n't". This ensures that the downstream model clearly recognizes the presence of negation.

The $ and ! symbols are isolated as independent tokens, allowing the next data-cleaning step to easily remove them or convert them into entity labels with financial meaning.

3.2. Complex preprocessing for Vietnamese with Underthesea

image.png

Now turn to a language with a completely different morphological character: Vietnamese. Vietnamese is an isolating language, where whitespace does not mark the boundary of a "word" (word / Linguistically Meaningful Unit) but only separates "syllables." If you blindly apply English NLTK methods to Vietnamese, tightly bound compound words such as "thời gian" will be shattered into two meaningless fragments, "thời" and "gian".

To thoroughly solve this barrier, using Word Segmentation tools trained specifically on Vietnamese corpora is a mandatory preprocessing step. Underthesea is a leading open-source library that provides a powerful API based on Conditional Random Fields (CRF) to bind discrete syllables into complete LMUs.

# Environment setup instructions: pip install underthesea

from underthesea import word_tokenize

def vietnamese_word_segmentation(text: str):
    """
    Demonstrate reconstructing syllables into semantically meaningful words in Vietnamese.
    This is a mandatory prerequisite before feeding text into any Transformer model.
    """
    print(f"--- RAW VIETNAMESE TEXT ---")
    print(f"{text}\n")
    
    # Perform word segmentation.
    # The format="text" parameter is designed to automatically join syllables belonging to the same word
    # with an underscore (_), forming a unified, inseparable block.
    tokens_with_underscore = word_tokenize(text, format="text")
    
    print(f"--- WORD SEGMENTATION RESULT ---")
    print(f"Output string: {tokens_with_underscore}")

vn_sample = "Tôi là sinh viên học tại trường đại học FPT Cần Thơ."
vietnamese_word_segmentation(vn_sample)

The text string will be transformed into: Tôi là sinh_viên học tại trường đại_học FPT Cần_Thơ.

By tightly binding syllables with _, core concepts such as "sinh_viên", "đại_học", and the proper noun "Cần_Thơ" preserve their semantic integrity. Thanks to this operation, later-stage BPE algorithms will treat "sinh_viên" as a unified block, avoiding incorrectly learned context. This technique has been proven to be a vital foundation for improving classification performance in practice.

3.3. Modern Subword Tokenization for LLMs with Hugging Face

The peak of modern Tokenization art lies in operating Subword structures through Hugging Face's transformers library. This is the gateway that feeds data into LLM models. The code below will implement the WordPiece system for an English BERT model and the BPE system for a Vietnamese PhoBERT model, thereby exposing their remarkable vocabulary-"shredding" mechanism.

# Environment setup instructions: pip install transformers torch

from transformers import AutoTokenizer

def modern_subword_tokenization():
    """
    Implement Subword-level Tokenization as used in Transformer architectures.
    Demonstrate Input ID generation and suffix-splitting behavior.
    """
    
    print("=== 1. BERT MODEL WORDPIECE SYSTEM (ENGLISH) ===\n")
    # Initialize the Tokenizer from the Hugging Face hub
    bert_tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
    en_text = "Transformers are revolutionizing natural language preprocessing!"
    
    # The tokenizer() function automatically performs every step: cleaning, subword splitting, and ID mapping
    bert_output = bert_tokenizer(en_text)
    
    # Convert the IDs back into string form to observe how the word is cut into pieces
    bert_tokens = bert_tokenizer.convert_ids_to_tokens(bert_output['input_ids'])
    
    print(f"Raw text: {en_text}")
    print(f"Subword list: {bert_tokens}")
    print(f"Input IDs vector:   {bert_output['input_ids']}\n")
    
    
    print("=== 2. PHOBERT MODEL BPE SYSTEM (VIETNAMESE) ===\n")
    # PhoBERT requires input to be word-segmented in advance (using Underthesea or RDRSegmenter)
    phobert_tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base")
    vn_text = "Tôi đang học về xử_lý ngôn_ngữ tự_nhiên."
    
    phobert_output = phobert_tokenizer(vn_text)
    phobert_tokens = phobert_tokenizer.convert_ids_to_tokens(phobert_output['input_ids'])
    
    print(f"Original text (already segmented): {vn_text}")
    print(f"Subword list:   {phobert_tokens}")
    print(f"Input IDs vector:     {phobert_output['input_ids']}")

modern_subword_tokenization()

With BERT: Learners will witness the magic of WordPiece. A long, complex word such as "revolutionizing" is neatly cut into two pieces: ["revolution", "##izing"]. The word "preprocessing" becomes ["pre", "##processing"]. The ## symbol is the system's indicator that this fragment is a continuation piece that must be joined to the root immediately before it to form a complete word. In addition, BERT automatically inserts structural navigation tokens such as [CLS] at the beginning of the sentence and [SEP] at the end to mark sequence boundaries.

With PhoBERT: The BPE algorithm splits the already-normalized Vietnamese string. The result shows compound words such as xử_lý and ngôn_ngữ encoded perfectly. The difference is that PhoBERT uses <s> and </s> tags to manage input structure instead of [CLS] and [SEP]. Mastering these differences in directional tokens is a survival skill for every NLP engineer.

4. The Art of Normalization and Preprocessing Strategy

Designing a Tokenization system does not stop at calling a few library commands. It is an art that interweaves digital thinking with deep linguistic knowledge. In a real-world data environment riddled with chaos from social media, forums, and unstructured text, a seasoned engineer must confront and resolve a host of noise traps before the data reaches the door of the neural network.

4.1. Text cleaning and normalization techniques (Text Normalization)

The data-cleaning process plays a decisive role in the purity of the vocabulary. Raw data from the Internet (Non-Standard Words - NSW) contains countless nonstandard variants capable of collapsing every deep-learning effort of an algorithm.

First is dealing with extra whitespace and digital noise. Texts are often stuffed with meaningless tab characters, blank lines, or extra spaces. Failing to clean them leads to an explosion of empty tokens, diluting information and unnecessarily increasing sequence length. The professional recommendation is to use libraries such as Clean-text or regular expressions (Regex) to flatten and normalize the text space. Likewise, hyperlinks (URLs), HTML tags, or email addresses rarely contribute semantic value but pollute the vocabulary severely. Engineers need to proactively filter them or replace them with a unique recognition token (for example: <URL>, <EMAIL>).

Next is the art of handling Uppercase / Lowercase (Casing). Inconsistency in distinguishing uppercase and lowercase causes massive duplication in identifier representation. For example, "Good", "good", and "GOOD" will be understood by the system as three completely separate concepts, dispersing the model's statistical strength and hindering its ability to understand sentiment. The decision here depends tightly on the task objective. If you are building a Named Entity Recognition (NER) system, capitalization is a vital clue for distinguishing the proper name "Apple" (the company) from the common noun "apple" (the fruit). However, for Sentiment Analysis, the wisest strategy is to convert the entire text to lowercase in order to compress vocabulary size and optimize learning without losing important information.

Another contemporary challenge is Emoji and special punctuation. In contemporary digital communication, a single emoji can carry more emotional expressive power than dozens of words combined. Using Regex to wipe out emoji is a fatal mistake in customer-feedback analysis. Instead, experts recommend two solutions: ensure the Tokenizer has a mechanism to decode Unicode so that emoji is protected as an independent, high-value token; or use translators (such as the emoji library in Python) to convert them into text form (for example, an angry-face emoji into [angry_face]) so the model can absorb them easily. Likewise, for punctuation, a sequence of periods ... or exclamation marks !!! is a clear signal of sarcasm or escalating emotion, and they need to be tokenized into independent semantic units rather than glued to the preceding word.

In Vietnam, organizations and engineers often need to leverage specialized local normalization tools. Libraries such as VietNormalizer provide the ability to transform integers, dates, currencies (VND, USD), and decode abbreviations into standard pronunciation form. This process not only shrinks the infinite variation of social-media language but also supplies clean, tightly structured data for both Text-to-Speech systems and large-language-model NLP.

4.2. Criteria for choosing a strategy for each language group

There is no Tokenizer algorithm that is universal ("one-size-fits-all"). Differences in morphological typology among language groups directly determine the choice of tool. Forcing a tool into the wrong role will result in a broken machine-learning architecture.

Languages with whitespace-delimited words (English, Indo-European languages):

Because word boundaries are already marked transparently by spaces, experts can directly deploy Subword models (BPE, WordPiece, SentencePiece) on raw data. Even so, a preprocessing layer that intelligently splits contractions (such as "don't" or "I'm") with NLTK or SpaCy before feeding them into a subword tokenizer remains a standard practice for keeping the data pure.

Languages written without whitespace (Chinese, Japanese, Korean):

Whitespace is completely absent from the structure of these languages, rendering any space-based tokenization effort meaningless. The solution here is to tightly combine complex Morphological Analyzers that integrate linguistic dictionaries. These systems compute statistical probabilities to find the most reasonable word-cut boundaries. Another approach used strongly in modern LLMs is to fall back to Character-level Tokenization for these languages, accepting longer sequence lengths to ensure no semantic loss.

Isolating languages with syllable boundaries (Vietnamese):

Vietnamese has a unique characteristic: it is a monosyllabic language, yet most of its vocabulary consists of compound words. Whitespace exists but only separates syllables, not word boundaries (LMUs). Feeding raw Vietnamese text directly into a BPE model is an act of destroying linguistic structure.

The Ultimate Pipeline: An NLP engineer in Vietnam must strictly follow a two-stage process. In stage 1, the text must be processed by a machine-learning-based word segmentation tool. One can choose RDRSegmenter (based on the Ripple Down Rules algorithm, often recommended by the VnCoreNLP and PhoBERT systems), or Underthesea (using a CRF architecture that reaches up to 80% accuracy on standard datasets), or lightweight packages such as pyvi. This process welds related syllables together with an underscore (the phrase "chăm sóc" becomes "chăm_sóc"). Moving to stage 2, this already-structured data sequence is only then officially pushed into a Subword algorithm (such as PhoBERT's fastBPE). The combination of these two layers has been experimentally proven to be the core factor that delivers superior performance for every NLP model tackling Vietnamese tasks.

5. Case study: analyzing the practical impact of Tokenization

Abstract theory needs to be proven with real numbers. The two case studies below will expose the truth about how choosing the right or wrong Tokenization strategy can completely reverse the economics and performance of an Artificial Intelligence project.

Scenario 1: The tragedy of cost and performance when training a multilingual LLM

System context: In the global race to develop AI, a research consortium conducted large-scale training of 24 Transformer-architecture LLM models at a scale of 2.6 billion parameters, aiming to serve multiple languages. To accelerate progress and save design effort, the team decided to reuse a Tokenizer already optimized for English (an English-centric tokenizer) to train all other languages, hoping the BPE algorithm could automatically generalize and "learn" the structure of peripheral languages.

Course of events and consequences of the wrong choice: Forcing a Tokenizer that prioritizes English morphological structure onto distant language families ignited a comprehensive system disaster.

Downstream Performance Degradation: Languages with complex morphology or ones far from English (such as Vietnamese, Russian, or Greek) were brutally fragmented. Instead of recognizing a word with complete meaning, the English Tokenizer shredded the vocabulary of these languages into 5 to 7 discrete, meaningless character tokens. This rupture paralyzed the Attention mechanism inside the Transformer network, causing the model to completely lose its bearings when chaining context and performing logical inference.

Training Costs Explosion: The vocabulary-shredding phenomenon directly caused a sudden increase in token density per unit of information (a metric researchers call fertility). When a simple sentence consumes dozens of tokens to represent, the model's limited context length fills up quickly. The most horrific reported consequence was that computational cost surged by as much as 68%. Tens of thousands of processing hours on expensive GPU clusters were wasted merely digesting an irrational vocabulary produced by the preprocessing pipeline.

Macro-architectural solution: From this disaster, experts distilled a rule: a Tokenizer is not a neutral component. Surveys show that even designing a Tokenizer for the 5 largest European languages requires vocabulary size to swell to 3 times that of an English-only model in order to maintain parity. For low-resource languages such as Vietnamese or Greek, allocating budget to build a localized Tokenizer or using unlabeled, tokenization-free architectures (such as T-Free models) is a survival condition for closing the gap with giant technology corporations.

Scenario 2: A strategy for decoding customer-feedback sentiment in retail

System context: At a large Portuguese retail group expanding into the Vietnamese market, leadership wanted to apply Machine Learning to extract economic value from hundreds of thousands of customer comments on social media and e-commerce platforms. Accurate Sentiment Analysis would determine product-strategy adjustments, prevent communications crises, and position the market.

Experiment and a life-or-death decision:

Consider a typical high-noise customer comment: "Sản phẩm này giá không rẻ, nhưng dịch vụ chăm sóc khách hàng cực kỳ tốt."

The failure of traditional Word-level Tokenization: In the first attempt, the data team used a basic splitting method via the split() function and filtered out stopwords. The system tore the sentence into the array: *['Sản', 'phẩm', 'này', 'giá', 'không', 'rẻ', ',', 'nhưng', 'dịch', 'vụ', 'chăm', 'sóc', 'khách', 'hàng', 'cực', 'kỳ', 'tốt']*. The Bag-of-Words or TF-IDF algorithm that followed received the words "rẻ" and "tốt" as isolated entities, while the vital negation "không" risked being discarded or failing to link with "rẻ". The result? A naive classification model judged the entire comment as Positive for both pricing policy and service, pushing the pricing department into a dangerous strategic trap.

The rise of a combined preprocessing architecture (Segmentation + BERT): Recognizing the fatal weakness, they applied a sophisticated Tokenization strategy: start by cleaning and normalizing the text, then use the Underthesea word-segmentation tool. Phrases were merged into LMUs with tight semantic structure: không_rẻ, chăm_sóc, khách_hàng, cực_kỳ. Finally, this array was pushed through BERT's WordPiece algorithm. At that point, BERT's Attention mechanism, combined with data that had been correctly grouped into semantic blocks, easily identified and isolated the Negative tone of price (không_rẻ) and the Positive tone of service.

Economic and technical outcome: Designing a Tokenization strategy that respects the morphological structure of the data completely transformed the system. In an evaluation study of a similar algorithm on a product-review dataset, a BERT architecture combined with an optimized vocabulary crushed classical models (such as Logistic Regression or SVM), establishing an extraordinary Accuracy of 94.2% and an AUC-ROC approaching perfection at 0.97. Looking deeper into the component metrics, well-tokenized systems delivered a sharp balance between Precision and Recall, pushing the F1 Score to a confidence threshold that allowed complete automation of the customer-service flow.

Through two practical slices — at the level of core model training and at the level of commercial application — the full picture has been drawn vividly. Splitting a piece of text may at first glance look like a modest preprocessing operation, but the ripple effects it creates have the power to determine the sophistication of a cognitive system, decide the efficiency of allocating millions of dollars in server infrastructure cost, and shape a company's core competitive capability in the data marketplace. Mastering the art of Tokenization is not merely understanding a snippet of programming code; it is the key that unlocks a deep understanding of how humanity communicates and how artificial intelligence thinks.

References

Summary of the tokenizers, https://huggingface.co/docs/transformers/v4.44.1/en/tokenizer_summary

Huỳnh Phước Nguyên

Written by Huỳnh Phước Nguyên

AI Engineer, BK Hightech

Ready to build something great?

Tell us about your project and we'll get back to you within a day.

Get in Touch