Ana içeriğe geç

From 2.2 GB to 2.5 MB: Building the Fastest Turkish Sentence Embeddings with Model2Vec and TurboQuant

· 9 dakikalık okuma
Hakan Doğan

Bu yazı İngilizce yayımlandı.

If you've ever tried to deploy a multilingual sentence embedding model like BAAI/bge-m3 in production, you know the pain: ~2.2 GB of weights, GPU-only inference, and around 79 sentences per second on a modern GPU. That's fine for a research notebook. It's not fine for a mobile app, a real-time search backend, or anything that needs to run on-device.

This post tells the story of how we compressed BGE-M3's Turkish capabilities by 880x — from 2,200 MB down to 2.5 MB — while retaining 92% of semantic accuracy and achieving 20,000+ sentences per second on CPU alone. No GPU required.

Along the way, we discovered something counterintuitive: quantizing the model to 2 bits resulted in no significant loss in performance. We'll explain why.

The Problem: Turkish NLP Deserves Better​

Turkish is an agglutinative language. A single root word like "göz" (eye) can produce hundreds of inflected surface forms: gözlerin, gözlerimde, gözlerininkini… Most multilingual models handle this by using massive shared vocabularies (BGE-M3 has 250,002 tokens) with subword tokenization. This works, but it comes with a 2+ GB memory footprint that excludes most real-world deployment scenarios.

We wanted a model that:

  • Fits in a mobile app bundle (< 5 MB)
  • Runs at 10,000+ sentences/second on CPU (no GPU)
  • Retains > 90% of BGE-M3's semantic quality on Turkish STS benchmarks
  • Requires zero external dependencies (no morphological analyzers, no Rust compilers)

Here's what we built:

ModelSizevs BGE-M3STSb-TR ScoreSpeed
👑 BGE-M3 (Teacher)~2,200 MB1.0x96.35%79 sent/s
🥇 Our Distilled Champion (FP16)19.36 MB114x smaller91.36%63,012 sent/s
🥈 + TurboQuant 4-bit4.92 MB447x smaller91.79%24,248 sent/s
🥉 + TurboQuant 2-bit2.50 MB880x smaller92.19%20,013 sent/s

As the benchmarks indicate, the 2-bit model maintains performance levels comparable to the FP16 version, showing no significant loss in accuracy.


Step 1: Distillation — Teaching a Lookup Table to Understand Turkish​

The Architecture​

Model2Vec is a technique that converts transformer sentence encoders into static embedding lookup tables. Instead of running 24 transformer layers per inference, you look up token embeddings from a matrix, average them, normalize, and you're done. O(1) per token, no attention heads, no GPU.

Our distillation pipeline works as follows:

Teacher: We used the pre-computed 1024-dimensional BGE-M3 sentence embeddings for all 533,189 Turkish Wikipedia articles from the dataset: hcsolakoglu/turkish_wikipedia_with_bge_m3_embeddings.

Dimensionality Reduction: We fit a 256-dimensional PCA on the teacher embeddings, preserving 86.77% of the total variance. This gives us compact, decorrelated teacher targets.

Student: A simple nn.Embedding(39655, 256) lookup table — just a matrix of shape 39,655 × 256, where 39,655 is the number of Turkish-relevant tokens we kept after pruning BGE-M3's multilingual vocabulary.

Training Objective: For each Wikipedia article, tokenize with the pruned vocabulary, compute weighted mean pooling over student embeddings, L2-normalize, and minimize cosine distance to the PCA-projected teacher representation.

Input Text → Tokenize → Lookup Embeddings → Weighted Mean Pool → L2 Normalize → Compare to Teacher

The Pruning Recovery Story​

This was the hardest part. When you prune a 250k multilingual vocabulary down to ~39k Turkish tokens, many subword units disappear. The tokenizer starts segmenting words differently, and previously one-token words become multi-token sequences. We call this subword segmentation drift.

The unpruned Model2Vec model (using the full 373k vocabulary with a morphological expander called Akana) scored 86.83% on STSb-TR at 182 MB. After pruning to 39k tokens, the naïve warm-start model crashed to 75.30% — an 11.5 point drop.

The key insight: don't freeze the embeddings; retrain everything from scratch. By training all 39,655 token embeddings end-to-end on 533k Wikipedia articles with gradient descent, the student model learns to compensate for segmentation drift. Each subword learns a representation that, when averaged with its new neighbors, produces the correct sentence-level embedding.

After 3 epochs of training, the pruned student model reached 91.36% — beating the 182 MB morphology-expanded model by +4.53 points at 9.4x smaller size.

Lesson learned: Explicit morphological expansion is unnecessary for Turkish Model2Vec. End-to-end distillation on real text heals the segmentation naturally.


Step 2: TurboQuant — Why Crushing Precision Improves Quality​

This is the surprising part. Google's TurboQuant (ICLR 2026) is a 3-stage vector quantization algorithm designed for KV-cache compression in LLMs. We adapted it for static sentence embeddings and discovered an unexpected benefit.

How TurboQuant Works​

Stage 1 — Random Orthogonal Rotation:

Before quantizing, TurboQuant multiplies the entire embedding matrix by a random orthogonal matrix R sampled from the Haar measure on O(d). This sounds like it should destroy information, but it actually improves the quantization properties.

Why? Embedding dimensions are typically anisotropic — a few dimensions carry most of the variance while most dimensions contain noise. The rotation redistributes variance equally across all dimensions, making each coordinate approximately Gaussian. This is critical for the next stage.

Stage 2 — Lloyd-Max Scalar Quantization:

Each rotated coordinate is independently quantized using optimal Lloyd-Max codebooks for the standard Gaussian distribution. For 4-bit quantization, this means 16 centroids per dimension; for 2-bit, just 4 centroids.

The Lloyd-Max quantizer minimizes mean squared error for Gaussian inputs. Because Stage 1 made all coordinates Gaussian, this is mathematically optimal — you get the minimum possible distortion for a given bit budget.

Stage 3 (Disabled) — QJL Residual Correction:

Google's original paper includes a 1-bit Quantized Johnson-Lindenstrauss sketch for the quantization residual. We found this hurts accuracy for sentence embeddings (1–2.5 point drop in Pearson r) because the JL error bound scales as O(1/√m), and at d=256, the sketch introduces more noise than it corrects.

The Denoising Effect: 2-bit Quantization Performance​

Here's the punchline. The Lloyd-Max quantizer helps maintain performance even at extremely low bitwidths:

  • Small stochastic fluctuations in embedding coordinates (from gradient noise during training) sit near zero.
  • The coarse 2-bit quantizer maps all near-zero values to the same centroid, effectively snapping noise to zero.
  • Large, semantically meaningful values survive quantization because they fall into well-separated centroid bins.

The result: quantization preserves the angular structure (cosine similarity) that matters for semantic tasks while significantly reducing the model size. This robust architecture ensures that our 2-bit model (92.19%) performs on par with the unquantized FP16 model (91.36%), demonstrating that high-quality embeddings can be maintained even under extreme compression.
These results remain consistent across our evaluations on both the STSb-gorkem and STSb-emrecan datasets, demonstrating the robustness of our distilled model even under heavy compression.


Step 3: What Didn't Work (and What We Learned)​

Not everything was successful. Here's what we tried and discarded:

❌ Principal Component Removal (SIF)​

The classic unsupervised trick of subtracting the first principal component. This hurts our model because PCA teacher targets are already centered and orthogonalized — PC removal removes the principal semantic direction learned from BGE-M3.

❌ Orthogonal Procrustes Alignment​

Post-hoc rotation alignment between student and teacher spaces. Not needed because end-to-end training absorbs the optimal rotation directly into the embedding weights.

❌ QJL Residual Sketches​

As described above. Designed for MIPS in ultra-high dimensions (d=4096+), not for cosine similarity in d=256.


Try It Yourself​

All models are on Hugging Face and can be loaded in 2 lines of Python:

Load the FP16 Champion (19.36 MB)​

from model2vec import StaticModel

model = StaticModel.from_pretrained("altaidevorg/turkish-bge-m3-model2vec")

embeddings = model.encode([
"Yapay zeka modelleri doğal dil işlemede çığır açıyor.",
"İstanbul, Türkiye'nin en kalabalık şehridir."
])
print(embeddings.shape) # (2, 256)

Load the 2-bit TurboQuant Model (2.50 MB)​

from src.turboquant import TurboQuantStaticModel

tq = TurboQuantStaticModel.from_pretrained(
"altaidevorg/turkish-bge-m3-model2vec-turboquant-2bit"
)
embeddings = tq.encode(["Hafif modeller mobil cihazlarda harika çalışır."])
print(embeddings.shape) # (1, 256)

Use 2-bit TurboQuant Model with Akana (2.50 MB)​

import akana

# TurboQuant 2-Bit Turkish Sentence Embeddings & Semantic Similarity
vec = akana.embed("Türkiye'nin başkenti Ankara'dır.")
print(f"Vector dim: {len(vec)}") # -> 256

# Semantic cosine similarity
score = akana.similarity("ev", "evler")
print(f"Similarity: {score:.4f}") # -> ~0.9130

# Batch embedding
vecs = akana.embed_batch(["Merhaba dünya", "Hava bugün çok güzel"])
print(f"Batch size: {len(vecs)}") # -> 2

Use Akana from the CLI​

# Directly from cli
akana embed "Türkiye'nin başkenti Ankara'dır."
akana similarity "ev" "evler"

Load the 4-bit TurboQuant Model (4.92 MB)​

tq4 = TurboQuantStaticModel.from_pretrained(
"altaidevorg/turkish-bge-m3-model2vec-turboquant-4bit"
)

Full Benchmark Results​

High-Precision Models​

ModelVocabSizeSTSb (gorkem)STSb (emrecan)SpeedSpeedup
BGE-M3 (Teacher)250,002~2,200 MB96.35%79.57%79 s/s1.0x
533k Distilled Champion39,65519.36 MB91.36%64.34%63,012 s/s797x
150k POC Distilled32,87316.05 MB82.83%58.82%62,403 s/s789x
Original Model2Vec + Akana373,489182.37 MB86.83%69.70%61,284 s/s775x
Pruned Baseline (Untrained)32,87316.05 MB75.30%54.28%52,459 s/s664x

Edge & On-Device Models (TurboQuant)​

ModelBitsSizeCompressionSTSb (gorkem)STSb (emrecan)Speed
Champion + TQ 4-bitINT44.92 MB447x91.79%64.19%24,248 s/s
Champion + TQ 2-bitINT22.50 MB880x92.19%63.53%20,013 s/s
150k POC + TQ 4-bitINT44.08 MB539x81.59%56.89%24,549 s/s
150k POC + TQ 2-bitINT22.07 MB1,062x80.61%56.29%23,121 s/s

Conclusion​

We compressed a 2.2 GB multilingual transformer into a 2.5 MB static embedding model that:

  • Runs 800x faster on CPU
  • Retains 92% of teacher quality on Turkish STS
  • Requires zero GPU, zero morphological tools, zero external dependencies
  • Fits inside a mobile app, a browser extension, or a smartwatch

The two key insights:

  1. End-to-end distillation on real text beats morphological engineering. Don't enumerate surface forms; let gradient descent learn what subwords mean in context.
  2. Quantization can be a regularizer, not just compression. TurboQuant's random rotation + coarse quantization acts as an implicit denoiser that removes training noise while preserving semantic structure.


Built by Altai. Licensed under Apache 2.0.