LLM Practitioner’s Guide: Falcon, Mistral, Smaug
So, you want to fine-tune a large language model for your own language. What model should you use?
The world of open-access LLMs is rapidly expanding, with hardly a week passing without the announcement of a new benchmark leader. At the time of writing this post, the Huggingface Open LLM Leaderboard has crowned a new champion: Smaug LLM. This field is incredibly dynamic, reminding us of the bustling IBM PC clone market in the late 1980s to early 1990s.
Despite the variety, many of these models share common origins. Smaug, for example, is based on Qwen-72B. While fine-tuning allows for significant adjustments to a model\u2019s behavior, one aspect you can\u2019t alter is its tokenizer.

We’ve been deeply involved with customizing, fine-tuning, and deploying various open-access LLMs. Through this series of posts, we aim to share our insights on what they can and can’t do, along with detailed instructions and best practices for fine-tuning.
LLM Practitioner’s Guide:
- How Multilingual Llama-2 Actually Is?
- How Multilingual Falcon, Mistral, Smaug, and other LLMs Are?
In a previous post, we explored the Llama-2 tokenizer and discussed why it struggles with certain languages. Today, we\u2019re conducting comprehensive tokenizer tests on six of the most popular LLMs across 12 languages. Our goal is to determine which models are the most adaptable for each language. We\u2019ll also share the source code for our experiments, enabling you to include any language you wish in this benchmark.
All large language models (LLMs), whether proprietary like ChatGPT or open-source, process information using numbers. To interpret text, these models initially convert it into numerical form. The main challenge in natural language processing (NLP) lies in finding the most efficient method to represent languages as numbers.
A simple approach assigns each letter or symbol a unique number. The model then attempts to predict the next number in a sequence. However, this approach can turn a brief sentence into a lengthy sequence of 61 numbers. Any error in predicting a single letter can affect all subsequent predictions, making this method both complex and prone to mistakes.
To address these challenges, researchers have developed a method known as tokenization. Tokenization breaks down words into smaller units, or tokens, that are common across various words. This approach means the model doesn’t have to predict each letter or entire word but rather smaller, more manageable sequences of tokens. By significantly reducing the number of options the model needs to consider—from an entire dictionary to roughly 2,000 tokens for English—it makes the process both more efficient and accurate.
Test Setup: Six Models, Twelve Languages
Let’s examine how today’s most popular LLM tokenizers handle various languages. We will be testing the following models:
- Mixtral-8x7B-Instruct-v0.1
- Mistral-7B-Instruct-v0.2
- falcon-7b-instruct
- Llama-2-7b-chat-hf
- Smaug-72B-v0.1
- mt5
We included the mt5 LLM in our lineup because it boasts one of the best multilingual tokenizers, making it an excellent benchmark for comparing the other models on our list.
First, we need to import the tokenizers:
from transformers import AutoTokenizer
tokenizers = {
'Mixtral-8x7B-Instruct-v0.1': AutoTokenizer.from_pretrained("mistralai/Mixtral-8x7B-Instruct-v0.1"),
'Mistral-7B-Instruct-v0.2': AutoTokenizer.from_pretrained("mistralai/Mistral-7B-Instruct-v0.2"),
'falcon-7b-instruct': AutoTokenizer.from_pretrained("tiiuae/falcon-7b-instruct"),
'Llama-2-7b-chat-hf': AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-chat-hf"),
'Smaug-72B-v0.1': AutoTokenizer.from_pretrained("abacusai/Smaug-72B-v0.1"),
'mt5': AutoTokenizer.from_pretrained("google/mt5-small")
}Next, we’ll select languages and their well-known sentences for testing. If you’re interested in adding another language to the test, simply include it as a new entry in the dictionary below.
sentences = {
'English': 'Call me Ishmael. Some years ago...',
'German': 'Da steh ich nun, ich armer Tor!...',
'French': "On ne voit bien qu'avec le cœur...",
'Spanish': "En un lugar de la Mancha...",
'Polish': 'Litwo, ojczyzno moja!...',
'Ukrainian': 'Як умру, то поховайте мене...',
'Greek': 'ἄνδρα μοι ἔννεπε, Μοῦσα...',
'Hebrew': 'בְּרֵאשִׁית בָּרָא אֱלֹהִים...',
'Arabic': "حدثني أبي، عن جدي...",
'Hindi': "कर्मण्येवाधिकारस्ते...",
'Chinese': "道可道,非常道。名可名,非常名。",
'Japanese': "私は、その男の写真を三葉..."
}Next, let’s define the testing process:
def tokenizer_test(sentences):
results = {}
for model_name, tokenizer in tokenizers.items():
results[model_name] = {}
for language, sentence in sentences.items():
token_list = tokenizer.tokenize(sentence)
results[model_name][language] = {
'length': len(token_list),
'tokens': token_list
}
return results
results = tokenizer_test(sentences)Results: Token Length by Model and Language
It’s important to note that a smaller number of tokens generally indicates a more efficient tokenizer: given the same input, a shorter list of tokens suggests a better understanding of the context. For LLMs, working with shorter sequences is also more manageable than dealing with longer ones.
| Language | Mixtral | Mistral | Falcon | Llama-2 | Smaug | mt5 |
|---|---|---|---|---|---|---|
| English | 54 | 54 | 53 | 54 | 52 | 59 |
| German | 21 | 21 | 21 | 21 | 21 | 20 |
| French | 25 | 25 | 21 | 23 | 23 | 24 |
| Spanish | 20 | 20 | 17 | 20 | 18 | 18 |
| Polish | 48 | 48 | 39 | 47 | 42 | 39 |
| Ukrainian | 32 | 32 | 79 | 29 | 34 | 23 |
| Greek | 53 | 53 | 60 | 53 | 48 | 26 |
| Hebrew | 59 | 59 | 96 | 59 | 72 | 35 |
| Arabic | 29 | 29 | 33 | 29 | 13 | 12 |
| Hindi | 81 | 81 | 127 | 85 | 75 | 34 |
| Chinese | 17 | 17 | 18 | 17 | 14 | 15 |
| Japanese | 22 | 22 | 26 | 24 | 14 | 11 |
Lower values = better tokenization. Highlighted values indicate the best tokenizer for that language.

Vocabulary Sizes
Each tokenizer comes with a different vocabulary size, which influences the amount of data required for fine-tuning the LLM. Let’s compare these vocabulary sizes.
vocab_sizes = {name: tokenizer.vocab_size for name, tokenizer in tokenizers.items()}
plt.figure(figsize=(10, 6))
plt.bar(range(len(vocab_sizes)), list(vocab_sizes.values()), tick_label=list(vocab_sizes.keys()))
plt.ylabel('Vocabulary Size')
plt.title('Vocabulary Sizes of Various Tokenizers')
plt.show()
Conclusions
The remarkable quality of mt5’s tokenization can largely be attributed to its extensive vocabulary size. However, the effectiveness of a tokenizer involves more than just its size. For instance, even though Smaug’s vocabulary size is significantly larger than that of Llama-2, their tokenization performance is relatively similar.
It’s clear that none of the popular LLMs match the efficiency of the mt5 tokenizer, with falcon-7b-instruct showing particular inefficiency on average. However, for certain languages, especially those within the Romance and Germanic families, the models perform comparably well. Yet, when we explore languages outside these families, the performance of popular LLMs in multilingual tasks drops significantly. An inferior multilingual tokenizer renders fine-tuning for such tasks nearly impossible.
In our projects, we attempted to fine-tune Llama-2 for multilingual tasks using a vast and carefully curated dataset of over a billion tokens, only to switch to mt5 partway through the project.
Since mt5 only comes as a base model and requires substantial effort to teach it to follow instructions, it necessitates a large and high-quality dataset. Our clients were fortunate to have such a resource, but overall it is a rare and publicly unavailable asset.
Despite the impressive landscape of LLMs, there remains a gap for a high-performing multilingual model. The design of tokenizers in the most powerful open-access LLMs inherently limits their effectiveness in certain languages, and this is not something that can be overcome with additional training. Currently, training mt5 from the base model appears to be the most promising approach for achieving a truly multilingual model. Without it, creating an open-access LLM that is genuinely multilingual remains a challenge next to impossible.
Building your custom LLM could enhance data security and compliance and enable an AI competitive advantage for your product. Check our other posts: what the network effect is and how AI enables it, how to build an AI competitive advantage, what culture helps, what to avoid, and more.
Building AI for Non-English Markets?
Tokenizer design fundamentally limits most LLMs' multilingual capabilities. Let's identify the right model and training strategy for your language needs.