Shelpuk & Company
    Back to Blog
    February 26, 202412 min readSergii Shelpuk

    LLM Practitioner’s Guide: Multilingual Gemma

    Gemma, a Game-Changing Multilingual LLM

    LLM Practitioner’s Guide:

    On February 21, 2024, Google unveiled its new open-access Large Language Model (LLM), Gemma. This release is a significant milestone, much more important than it might initially seem.

    Ever since ChatGPT burst onto the scene, the realm of open-access LLMs has seen explosive growth. From the start, we’ve been deeply involved in training and customizing open-access chat LLMs. Despite the surrounding excitement, those of us working with multilingual models have been acutely aware of a critical issue. For languages outside the Romance and Germanic families, inefficient tokenizers have severely limited the performance of open-access LLMs.

    What’s a tokenizer, you might ask? It’s an algorithm that breaks down text into smaller pieces, known as tokens. These tokens can be words, parts of words, individual characters, or other units, depending on the tokenizer’s design. Tokenization is a crucial step in processing text for LLMs, enabling applications like text classification, question answering, and machine translation.

    Tokenizers are vital to the natural language processing (NLP) pipeline, serving a singular goal: converting text into a format that the model can understand. Since models operate on numerical data, tokenizers transform our textual inputs into numbers that the model can process.

    One simple approach is to assign a unique number to each letter or symbol. For example, “a” might be assigned the number 1, “b” the number 2, and so on. The model then tries to predict the next number in a sequence. However, this method can turn even a short sentence into a lengthy sequence of numbers. A single error in predicting a letter can throw off all subsequent predictions, making this method both complicated and prone to mistakes.

    An alternative strategy assigns each word a unique number. This condenses the sequence significantly, but it’s not without challenges. With more than 200,000 words in the English language alone, a model that works across multiple languages may need to differentiate among millions of numbers, significantly increasing the chance of errors.

    To address these challenges, researchers devised tokenization. It breaks words into smaller, commonly found pieces (tokens). For example, “predicting” might be split into “pred,” “ict,” and “ing.” This dramatically reduces the choices the model needs to consider—from an entire dictionary to around 2,000 tokens for English.

    However, there’s a significant issue: all the well-known open-access chat LLMs, such as Llama, Llama-2, Mistral, Mixtral, Falcon, Smaug, and others, have tokenizers primarily designed for Romance and Germanic languages. For non-Germanic, non-Romance, and non-Slavic languages such as Greek, Chinese, Hebrew, Arabic, Hindi, Korean, Japanese, etc.—these languages are reduced to characters.

    While you can enhance a model’s performance by training it with more or higher-quality data, the tokenizer presents a challenge that cannot be similarly addressed. The tokenizer operates based on a deterministic algorithm, not machine learning. You cannot train or modify it. The model is intricately linked to its tokenizer; attempting to replace it would make the model non-functional.

    There have been models with effective multilingual tokenizers, such as Google’s mt5, released in 2020. However, these were “base” models and not specifically trained for conversational tasks.

    This situation effectively limited the practical applications of LLMs to primarily Romance and Germanic languages. To achieve a high-quality conversational LLM, you had two choices:

    1. Fine-tune modern LLMs, such as Llama-2 or Mistral, in a way reminiscent of 2015, utilizing character-based text representation.
    2. Start with a base model equipped with a robust tokenizer and train it to be conversational.

    Both options demand extensive and highly accurate datasets. For most languages and companies, this made the fine-tuning of modern multilingual conversational LLMs practically out of reach.

    Until now.

    Gemma

    Gemma represents a new generation of lightweight, cutting-edge open models that build upon the research and technology behind the Gemini models. Developed by Google DeepMind and various Google teams, Gemma draws inspiration from Gemini, with its name stemming from the Latin word “gemma,” meaning “precious stone.”

    Here’s what you need to know:

    • Google has unveiled the Gemma model weights in two versions: Gemma 2B and Gemma 7B. Each version comes in pre-trained and instruction-tuned variants.
    • The newly introduced Responsible Generative AI Toolkit offers guidance and vital tools for developing safer AI applications using Gemma.
    • Google supports toolchains for inference and supervised fine-tuning (SFT) across all major frameworks, including JAX, PyTorch, and TensorFlow, natively supported in Keras 3.0.
    • Google has made available ready-to-use Colab and Kaggle notebooks, as well as integration with widely-used tools like Hugging Face, MaxText, NVIDIA NeMo, and TensorRT-LLM.
    • The pre-trained and instruction-tuned Gemma models run smoothly on various platforms, from your personal laptop to Google Cloud, with deployment options on Vertex AI and Google Kubernetes Engine (GKE).
    • The terms of use allow for responsible commercial use and distribution by organizations of any size.

    We previously highlighted mt5, Google’s 2020 LLM renowned for its exceptional multilingual tokenizer. Now, meet Gemma: equipped with an even more advanced tokenizer. In our latest analysis, we conducted thorough tests on the tokenizers of major LLMs across a variety of languages, and here’s how Gemma’s tokenizer stands out:

    Token length by model and language – Gemma comparison

    And here is a direct comparison of tokenization performance between Gemma and Llama-2:

    LanguageLlama-2 TokensGemma TokensImprovement
    English54507%
    German21205%
    French231917%
    Spanish201525%
    Polish473232%
    Ukrainian292417%
    Greek532847%
    Hebrew593934%
    Arabic291259%
    Hindi854053%
    Chinese171418%
    Japanese241250%

    At first glance, Gemma might seem like just another LLM – similar to the many others released by the open-source community each month. However, it’s important not to be misled; this is huge news for the LLM field. With Gemma’s introduction, companies worldwide now have the ability to fine-tune open-access LLMs for a variety of applications. This includes creating conversational translators, multilingual writing assistants, and Retrieval-Augmented Generation (RAG) systems. This release effectively bridges a significant, albeit previously overlooked, gap in the LLM landscape. It marks the arrival of the first truly multilingual chat LLM.

    Building your custom LLM could enhance data security and compliance and enable an AI competitive advantage for your product. Check our other posts for extensive explanations:

    Building a Multilingual AI Product?

    Choosing the right model and tokenizer is critical. Let's find the best foundation for your multilingual AI solution.

    Frequently Asked Questions

    Answers to common questions about this topic.

    How good is Gemma for multilingual LLM use cases?

    Gemma marks the arrival of the first truly multilingual chat LLM. Why this Google release is a game-changer for multilingual AI applications. Shelpuk AI Technology Consulting implements these patterns in production-focused engineering and AI delivery engagements.

    What does a team need to implement this solution reliably?

    A reliable implementation stack combines components such as LLM, Gemma, Multilingual, Tokenizer and aligns them in a clear delivery workflow. Shelpuk AI Technology Consulting supports teams with consulting and hands-on implementation of this stack.

    What performance or quality gains are realistic?

    Concrete technical gains are linked to specific architecture and delivery decisions. Shelpuk AI Technology Consulting helps organizations convert benchmark gains into durable production performance.

    What can fail first in production and how should teams prevent it?

    Key risks include insufficient language or benchmark coverage, biased test sets, and incomplete validation plans before production use. Shelpuk AI Technology Consulting builds governance and quality controls that reduce avoidable delivery failures.

    How should teams run a low-risk pilot before scaling?

    Start with one narrow workflow tied to explicit KPI targets, validate results, and expand scope only after evidence. Shelpuk AI Technology Consulting can guide rollout from pilot validation to production hardening and scaling.