Multilingual Text Classification with Scikit-LLM and Multilingual Embeddings

Learn how to build a multilingual text classification pipeline using BGE-M3 embeddings, Ollama, and Scikit-LLM—no language-specific models required.

Multilingual Text Classification with Scikit-LLM and Multilingual Embeddings

In this article, you will learn how to build a multilingual text classification pipeline using multilingual large language model (LLM) embeddings and Scikit-learn, without training separate models for each language.

Topics we will cover include:

  • What multilingual LLM embeddings are and why they eliminate the need for language-specific models.
  • How to set up a free, local embedding pipeline using Ollama, BGE-M3, and Scikit-LLM.
  • How to train and evaluate a logistic regression classifier on top of multilingual embeddings using a real-world review dataset.

Multilingual Text Classification with Scikit-LLM and Multilingual Embeddings

Introduction

Building machine learning models for a global audience, such as text classifiers based on multilingual data, traditionally required training a separate model for each language, making the process difficult to scale. Multilingual LLM embeddings solve this problem by producing numerical representations of text that map content from different languages into a shared vector space. With these language-agnostic embeddings, all that remains is training a lightweight downstream classifier on top of them. The following steps walk through how to do this using Scikit-LLM.

Initial Setup

The following commands install the required Python dependencies, system packages, and Ollama:

# Installing Python dependencies
pip install scikit-llm "datasets==2.19.1" -q

# Fix Colab's missing system dependencies first (version-dependent, use with care in other environments)
apt-get update -qq && apt-get install -y -qq zstd

# Installing Ollama distribution
curl -fsSL https://ollama.com/install.sh | sh

To keep the solution free and runnable across environments including notebooks, we use Ollama instead of paid APIs like OpenAI. Scikit-LLM will be configured to communicate with a local Ollama server running BGE-M3, a state-of-the-art open-source model that supports multilingual embedding generation.

Next, start the Ollama server as a background process — the most hassle-free approach for cloud-based notebooks — and pull the BGE-M3 model. More information about BGE-M3 is available on its official website.

import subprocess
import time

# Starting the Ollama server in the background
subprocess.Popen(["ollama", "serve"])
time.sleep(5)  # Give the server a few seconds to initialize

# Pulling the multilingual embedding model
ollama pull bge-m3

The final configuration step points Scikit-LLM to the local Ollama instance. This approach requires a dummy API key rather than a real one:

from skllm.config import SKLLMConfig

# Point Scikit-LLM to our local Ollama instance
SKLLMConfig.set_gpt_url("http://localhost:11434/v1/")

# Provide a dummy key (required by the internal client, but safely ignored by Ollama)
SKLLMConfig.set_openai_key("free-friendly-dummy-key")

Building the Pipeline

The first major step is loading the data. This pipeline uses the Amazon Multi-language Reviews dataset, which contains labeled customer reviews on a 5-star rating scale (internally encoded as labels 0 to 4). To keep execution time manageable — particularly during embedding generation — we load 1,000 English and 1,000 Spanish reviews for a total of 2,000 samples. You can use a larger sample if needed, but keep it language-balanced and ensure the data is randomly shuffled before applying a train-test split.

from datasets import load_dataset
import pandas as pd

print("Loading and shuffling data to ensure class diversity...")

# 1. Loading the complete split
# 2. Shuffling it randomly with shuffle()
# 3. Extracting 1000 varied samples with select(range(1000))
data_en = (load_dataset("mteb/amazon_reviews_multi", "en", split="train", trust_remote_code=True)
           .shuffle(seed=42)
           .select(range(1000)))

data_es = (load_dataset("mteb/amazon_reviews_multi", "es", split="train", trust_remote_code=True)
           .shuffle(seed=42)
           .select(range(1000)))

# Combining into a single DataFrame
df = pd.concat([pd.DataFrame(data_en), pd.DataFrame(data_es)], ignore_index=True)

# Shuffling bilingual data
df = df.sample(frac=1, random_state=42).reset_index(drop=True)

# Features and Labels
X = df['text']
y = df['label']

print(f"Total samples: {len(X)}")
print("\n--- Class Verification (should have samples from 0 to 4) ---")
print(y.value_counts())

Output:

Loading and shuffling data to ensure class diversity...
Total samples: 2000
--- Class Verification (should have samples from 0 to 4) ---

The dataset is now combined, shuffled, and ready for the embedding and classification steps. With multilingual embeddings handling the language-agnostic representation, a single logistic regression model trained on this mixed-language data can classify reviews in both English and Spanish without any language-specific preprocessing or model duplication.