Shelpuk & Company
    Back to Blog
    February 12, 202418 min readSergii Shelpuk

    LLM Practitioner’s Guide: Falcon, Mistral, Smaug

    So, you want to fine-tune a large language model for your own language. What model should you use?

    The world of open-access LLMs is rapidly expanding, with hardly a week passing without the announcement of a new benchmark leader. At the time of writing this post, the Huggingface Open LLM Leaderboard has crowned a new champion: Smaug LLM. This field is incredibly dynamic, reminding us of the bustling IBM PC clone market in the late 1980s to early 1990s.

    Despite the variety, many of these models share common origins. Smaug, for example, is based on Qwen-72B. While fine-tuning allows for significant adjustments to a model\u2019s behavior, one aspect you can\u2019t alter is its tokenizer.

    Multilingual LLMs Illustration

    We’ve been deeply involved with customizing, fine-tuning, and deploying various open-access LLMs. Through this series of posts, we aim to share our insights on what they can and can’t do, along with detailed instructions and best practices for fine-tuning.

    LLM Practitioner’s Guide:

    In a previous post, we explored the Llama-2 tokenizer and discussed why it struggles with certain languages. Today, we\u2019re conducting comprehensive tokenizer tests on six of the most popular LLMs across 12 languages. Our goal is to determine which models are the most adaptable for each language. We\u2019ll also share the source code for our experiments, enabling you to include any language you wish in this benchmark.

    All large language models (LLMs), whether proprietary like ChatGPT or open-source, process information using numbers. To interpret text, these models initially convert it into numerical form. The main challenge in natural language processing (NLP) lies in finding the most efficient method to represent languages as numbers.

    A simple approach assigns each letter or symbol a unique number. The model then attempts to predict the next number in a sequence. However, this approach can turn a brief sentence into a lengthy sequence of 61 numbers. Any error in predicting a single letter can affect all subsequent predictions, making this method both complex and prone to mistakes.

    To address these challenges, researchers have developed a method known as tokenization. Tokenization breaks down words into smaller units, or tokens, that are common across various words. This approach means the model doesn’t have to predict each letter or entire word but rather smaller, more manageable sequences of tokens. By significantly reducing the number of options the model needs to consider—from an entire dictionary to roughly 2,000 tokens for English—it makes the process both more efficient and accurate.

    Test Setup: Six Models, Twelve Languages

    Let’s examine how today’s most popular LLM tokenizers handle various languages. We will be testing the following models:

    • Mixtral-8x7B-Instruct-v0.1
    • Mistral-7B-Instruct-v0.2
    • falcon-7b-instruct
    • Llama-2-7b-chat-hf
    • Smaug-72B-v0.1
    • mt5

    We included the mt5 LLM in our lineup because it boasts one of the best multilingual tokenizers, making it an excellent benchmark for comparing the other models on our list.

    First, we need to import the tokenizers:

    from transformers import AutoTokenizer
    
    tokenizers = {
        'Mixtral-8x7B-Instruct-v0.1': AutoTokenizer.from_pretrained("mistralai/Mixtral-8x7B-Instruct-v0.1"),
        'Mistral-7B-Instruct-v0.2': AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2"),
        'falcon-7b-instruct': AutoTokenizer.from_pretrained("tiiuae/falcon-7b-instruct"),
        'Llama-2-7b-chat-hf': AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-chat-hf"),
        'Smaug-72B-v0.1': AutoTokenizer.from_pretrained("abacusai/Smaug-72B-v0.1"),
        'mt5': AutoTokenizer.from_pretrained("google/mt5-small")
    }

    Next, we’ll select languages and their well-known sentences for testing. If you’re interested in adding another language to the test, simply include it as a new entry in the dictionary below.

    sentences = {
        'English': 'Call me Ishmael. Some years ago...',
        'German': 'Da steh ich nun, ich armer Tor!...',
        'French': "On ne voit bien qu'avec le cœur...",
        'Spanish': "En un lugar de la Mancha...",
        'Polish': 'Litwo, ojczyzno moja!...',
        'Ukrainian': 'Як умру, то поховайте мене...',
        'Greek': 'ἄνδρα μοι ἔννεπε, Μοῦσα...',
        'Hebrew': 'בְּרֵאשִׁית בָּרָא אֱלֹהִים...',
        'Arabic': "حدثني أبي، عن جدي...",
        'Hindi': "कर्मण्येवाधिकारस्ते...",
        'Chinese': "道可道,非常道。名可名,非常名。",
        'Japanese': "私は、その男の写真を三葉..."
    }

    Next, let’s define the testing process:

    def tokenizer_test(sentences):
        results = {}
        for model_name, tokenizer in tokenizers.items():
            results[model_name] = {}
            for language, sentence in sentences.items():
                token_list = tokenizer.tokenize(sentence)
                results[model_name][language] = {
                    'length': len(token_list),
                    'tokens': token_list
                }
        return results
    
    results = tokenizer_test(sentences)

    Results: Token Length by Model and Language

    It’s important to note that a smaller number of tokens generally indicates a more efficient tokenizer: given the same input, a shorter list of tokens suggests a better understanding of the context. For LLMs, working with shorter sequences is also more manageable than dealing with longer ones.

    LanguageMixtralMistralFalconLlama-2Smaugmt5
    English545453545259
    German212121212120
    French252521232324
    Spanish202017201818
    Polish484839474239
    Ukrainian323279293423
    Greek535360534826
    Hebrew595996597235
    Arabic292933291312
    Hindi8181127857534
    Chinese171718171415
    Japanese222226241411

    Lower values = better tokenization. Highlighted values indicate the best tokenizer for that language.

    Tokenizer performance for Falcon, Mistral, Mixtral, Smaug, and other open-source LLMs

    Vocabulary Sizes

    Each tokenizer comes with a different vocabulary size, which influences the amount of data required for fine-tuning the LLM. Let’s compare these vocabulary sizes.

    vocab_sizes = {name: tokenizer.vocab_size for name, tokenizer in tokenizers.items()}
    
    plt.figure(figsize=(10, 6))
    plt.bar(range(len(vocab_sizes)), list(vocab_sizes.values()), tick_label=list(vocab_sizes.keys()))
    plt.ylabel('Vocabulary Size')
    plt.title('Vocabulary Sizes of Various Tokenizers')
    plt.show()
    Vocabulary sizes for Falcon, Mistral, Mixtral, Smaug, and other open-source LLMs

    Conclusions

    The remarkable quality of mt5’s tokenization can largely be attributed to its extensive vocabulary size. However, the effectiveness of a tokenizer involves more than just its size. For instance, even though Smaug’s vocabulary size is significantly larger than that of Llama-2, their tokenization performance is relatively similar.

    It’s clear that none of the popular LLMs match the efficiency of the mt5 tokenizer, with falcon-7b-instruct showing particular inefficiency on average. However, for certain languages, especially those within the Romance and Germanic families, the models perform comparably well. Yet, when we explore languages outside these families, the performance of popular LLMs in multilingual tasks drops significantly. An inferior multilingual tokenizer renders fine-tuning for such tasks nearly impossible.

    In our projects, we attempted to fine-tune Llama-2 for multilingual tasks using a vast and carefully curated dataset of over a billion tokens, only to switch to mt5 partway through the project.

    Since mt5 only comes as a base model and requires substantial effort to teach it to follow instructions, it necessitates a large and high-quality dataset. Our clients were fortunate to have such a resource, but overall it is a rare and publicly unavailable asset.

    Despite the impressive landscape of LLMs, there remains a gap for a high-performing multilingual model. The design of tokenizers in the most powerful open-access LLMs inherently limits their effectiveness in certain languages, and this is not something that can be overcome with additional training. Currently, training mt5 from the base model appears to be the most promising approach for achieving a truly multilingual model. Without it, creating an open-access LLM that is genuinely multilingual remains a challenge next to impossible.

    Building your custom LLM could enhance data security and compliance and enable an AI competitive advantage for your product. Check our other posts: what the network effect is and how AI enables it, how to build an AI competitive advantage, what culture helps, what to avoid, and more.

    Building AI for Non-English Markets?

    Tokenizer design fundamentally limits most LLMs' multilingual capabilities. Let's identify the right model and training strategy for your language needs.

    Frequently Asked Questions

    Answers to common questions about this topic.

    How multilingual are Falcon, Mistral, Smaug, and similar LLMs in practice?

    Comprehensive tokenizer tests on six popular LLMs across 12 languages. Which models are the most adaptable for each language. Shelpuk AI Technology Consulting implements these patterns in production-focused engineering and AI delivery engagements.

    What does a team need to implement this solution reliably?

    A reliable implementation stack combines components such as LLM, Multilingual, Tokenizer, NLP and aligns them in a clear delivery workflow. Shelpuk AI Technology Consulting supports teams with consulting and hands-on implementation of this stack.

    What performance or quality gains are realistic?

    Concrete technical gains are linked to specific architecture and delivery decisions. Shelpuk AI Technology Consulting helps organizations convert benchmark gains into durable production performance.

    What can fail first in production and how should teams prevent it?

    Key risks include insufficient language or benchmark coverage, biased test sets, and incomplete validation plans before production use. Shelpuk AI Technology Consulting builds governance and quality controls that reduce avoidable delivery failures.

    How should teams run a low-risk pilot before scaling?

    Start with one narrow workflow tied to explicit KPI targets, validate results, and expand scope only after evidence. Shelpuk AI Technology Consulting can guide rollout from pilot validation to production hardening and scaling.