LLM Practitioner’s Guide: Multilingual Gemma

LLM Practitioner’s Guide:
- How Multilingual Llama-2 Actually Is?
- How Multilingual Falcon, Mistral, Smaug, and other LLMs Are?
- Gemma, a Game-Changing Multilingual LLM
On February 21, 2024, Google unveiled its new open-access Large Language Model (LLM), Gemma. This release is a significant milestone, much more important than it might initially seem.
Ever since ChatGPT burst onto the scene, the realm of open-access LLMs has seen explosive growth. From the start, we’ve been deeply involved in training and customizing open-access chat LLMs. Despite the surrounding excitement, those of us working with multilingual models have been acutely aware of a critical issue. For languages outside the Romance and Germanic families, inefficient tokenizers have severely limited the performance of open-access LLMs.
What’s a tokenizer, you might ask? It’s an algorithm that breaks down text into smaller pieces, known as tokens. These tokens can be words, parts of words, individual characters, or other units, depending on the tokenizer’s design. Tokenization is a crucial step in processing text for LLMs, enabling applications like text classification, question answering, and machine translation.
Tokenizers are vital to the natural language processing (NLP) pipeline, serving a singular goal: converting text into a format that the model can understand. Since models operate on numerical data, tokenizers transform our textual inputs into numbers that the model can process.
One simple approach is to assign a unique number to each letter or symbol. For example, “a” might be assigned the number 1, “b” the number 2, and so on. The model then tries to predict the next number in a sequence. However, this method can turn even a short sentence into a lengthy sequence of numbers. A single error in predicting a letter can throw off all subsequent predictions, making this method both complicated and prone to mistakes.
An alternative strategy assigns each word a unique number. This condenses the sequence significantly, but it’s not without challenges. With more than 200,000 words in the English language alone, a model that works across multiple languages may need to differentiate among millions of numbers, significantly increasing the chance of errors.
To address these challenges, researchers devised tokenization. It breaks words into smaller, commonly found pieces (tokens). For example, “predicting” might be split into “pred,” “ict,” and “ing.” This dramatically reduces the choices the model needs to consider—from an entire dictionary to around 2,000 tokens for English.
However, there’s a significant issue: all the well-known open-access chat LLMs, such as Llama, Llama-2, Mistral, Mixtral, Falcon, Smaug, and others, have tokenizers primarily designed for Romance and Germanic languages. For non-Germanic, non-Romance, and non-Slavic languages such as Greek, Chinese, Hebrew, Arabic, Hindi, Korean, Japanese, etc.—these languages are reduced to characters.
While you can enhance a model’s performance by training it with more or higher-quality data, the tokenizer presents a challenge that cannot be similarly addressed. The tokenizer operates based on a deterministic algorithm, not machine learning. You cannot train or modify it. The model is intricately linked to its tokenizer; attempting to replace it would make the model non-functional.
There have been models with effective multilingual tokenizers, such as Google’s mt5, released in 2020. However, these were “base” models and not specifically trained for conversational tasks.
This situation effectively limited the practical applications of LLMs to primarily Romance and Germanic languages. To achieve a high-quality conversational LLM, you had two choices:
- Fine-tune modern LLMs, such as Llama-2 or Mistral, in a way reminiscent of 2015, utilizing character-based text representation.
- Start with a base model equipped with a robust tokenizer and train it to be conversational.
Both options demand extensive and highly accurate datasets. For most languages and companies, this made the fine-tuning of modern multilingual conversational LLMs practically out of reach.
Until now.
Gemma
Gemma represents a new generation of lightweight, cutting-edge open models that build upon the research and technology behind the Gemini models. Developed by Google DeepMind and various Google teams, Gemma draws inspiration from Gemini, with its name stemming from the Latin word “gemma,” meaning “precious stone.”
Here’s what you need to know:
- Google has unveiled the Gemma model weights in two versions: Gemma 2B and Gemma 7B. Each version comes in pre-trained and instruction-tuned variants.
- The newly introduced Responsible Generative AI Toolkit offers guidance and vital tools for developing safer AI applications using Gemma.
- Google supports toolchains for inference and supervised fine-tuning (SFT) across all major frameworks, including JAX, PyTorch, and TensorFlow, natively supported in Keras 3.0.
- Google has made available ready-to-use Colab and Kaggle notebooks, as well as integration with widely-used tools like Hugging Face, MaxText, NVIDIA NeMo, and TensorRT-LLM.
- The pre-trained and instruction-tuned Gemma models run smoothly on various platforms, from your personal laptop to Google Cloud, with deployment options on Vertex AI and Google Kubernetes Engine (GKE).
- The terms of use allow for responsible commercial use and distribution by organizations of any size.
We previously highlighted mt5, Google’s 2020 LLM renowned for its exceptional multilingual tokenizer. Now, meet Gemma: equipped with an even more advanced tokenizer. In our latest analysis, we conducted thorough tests on the tokenizers of major LLMs across a variety of languages, and here’s how Gemma’s tokenizer stands out:

And here is a direct comparison of tokenization performance between Gemma and Llama-2:
| Language | Llama-2 Tokens | Gemma Tokens | Improvement |
|---|---|---|---|
| English | 54 | 50 | 7% |
| German | 21 | 20 | 5% |
| French | 23 | 19 | 17% |
| Spanish | 20 | 15 | 25% |
| Polish | 47 | 32 | 32% |
| Ukrainian | 29 | 24 | 17% |
| Greek | 53 | 28 | 47% |
| Hebrew | 59 | 39 | 34% |
| Arabic | 29 | 12 | 59% |
| Hindi | 85 | 40 | 53% |
| Chinese | 17 | 14 | 18% |
| Japanese | 24 | 12 | 50% |
At first glance, Gemma might seem like just another LLM – similar to the many others released by the open-source community each month. However, it’s important not to be misled; this is huge news for the LLM field. With Gemma’s introduction, companies worldwide now have the ability to fine-tune open-access LLMs for a variety of applications. This includes creating conversational translators, multilingual writing assistants, and Retrieval-Augmented Generation (RAG) systems. This release effectively bridges a significant, albeit previously overlooked, gap in the LLM landscape. It marks the arrival of the first truly multilingual chat LLM.
Building your custom LLM could enhance data security and compliance and enable an AI competitive advantage for your product. Check our other posts for extensive explanations:
- What the network effect is
- How AI enables the network effect
- How to build an AI competitive advantage
- What culture helps build the right AI products
- What to avoid in your AI strategy
Building a Multilingual AI Product?
Choosing the right model and tokenizer is critical. Let's find the best foundation for your multilingual AI solution.