<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <id>https://altai.dev/tr/blog</id>
    <title>ALTAI Blog</title>
    <updated>2026-08-31T00:00:00.000Z</updated>
    <generator>https://github.com/jpmonette/feed</generator>
    <link rel="alternate" href="https://altai.dev/tr/blog"/>
    <subtitle>ALTAI Blog</subtitle>
    <icon>https://altai.dev/tr/favicon.svg</icon>
    <rights>© 2026 ALTAI</rights>
    <entry>
        <title type="html"><![CDATA[From 2.2 GB to 2.5 MB: Building the Fastest Turkish Sentence Embeddings with Model2Vec and TurboQuant]]></title>
        <id>https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant</id>
        <link href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant"/>
        <updated>2026-08-31T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[How we distilled a 24-layer multilingual transformer into a 2.5 MB static embedding model that runs 800x faster.]]></summary>
        <content type="html"><![CDATA[<p>If you've ever tried to deploy a multilingual sentence embedding model like <code>BAAI/bge-m3</code> in production, you know the pain: ~2.2 GB of weights, GPU-only inference, and around 79 sentences per second on a modern GPU. That's fine for a research notebook. It's not fine for a mobile app, a real-time search backend, or anything that needs to run on-device.</p>
<p>This post tells the story of how we compressed BGE-M3's Turkish capabilities by <strong>880x</strong> — from 2,200 MB down to <strong>2.5 MB</strong> — while retaining <strong>92% of semantic accuracy</strong> and achieving <strong>20,000+ sentences per second on CPU alone</strong>. No GPU required.</p>
<p>Along the way, we discovered something counterintuitive: <strong>quantizing the model to 2 bits resulted in no significant loss in performance</strong>. We'll explain why.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-problem-turkish-nlp-deserves-better">The Problem: Turkish NLP Deserves Better<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#the-problem-turkish-nlp-deserves-better" class="hash-link" aria-label="The Problem: Turkish NLP Deserves Better doğrudan bağlantı" title="The Problem: Turkish NLP Deserves Better doğrudan bağlantı" translate="no">​</a></h2>
<p>Turkish is an agglutinative language. A single root word like "göz" (eye) can produce hundreds of inflected surface forms: <em>gözlerin, gözlerimde, gözlerininkini…</em> Most multilingual models handle this by using massive shared vocabularies (BGE-M3 has 250,002 tokens) with subword tokenization. This works, but it comes with a 2+ GB memory footprint that excludes most real-world deployment scenarios.</p>
<p>We wanted a model that:</p>
<ul>
<li class="">Fits in a mobile app bundle (&lt; 5 MB)</li>
<li class="">Runs at <strong>10,000+ sentences/second on CPU</strong> (no GPU)</li>
<li class="">Retains <strong>&gt; 90% of BGE-M3's semantic quality</strong> on Turkish STS benchmarks</li>
<li class="">Requires <strong>zero external dependencies</strong> (no morphological analyzers, no Rust compilers)</li>
</ul>
<p>Here's what we built:</p>








































<table><thead><tr><th style="text-align:left">Model</th><th style="text-align:center">Size</th><th style="text-align:center">vs BGE-M3</th><th style="text-align:center">STSb-TR Score</th><th style="text-align:center">Speed</th></tr></thead><tbody><tr><td style="text-align:left">👑 BGE-M3 (Teacher)</td><td style="text-align:center">~2,200 MB</td><td style="text-align:center">1.0x</td><td style="text-align:center">96.35%</td><td style="text-align:center">79 sent/s</td></tr><tr><td style="text-align:left">🥇 Our Distilled Champion (FP16)</td><td style="text-align:center"><strong>19.36 MB</strong></td><td style="text-align:center"><strong>114x smaller</strong></td><td style="text-align:center"><strong>91.36%</strong></td><td style="text-align:center"><strong>63,012 sent/s</strong></td></tr><tr><td style="text-align:left">🥈 + TurboQuant 4-bit</td><td style="text-align:center"><strong>4.92 MB</strong></td><td style="text-align:center"><strong>447x smaller</strong></td><td style="text-align:center"><strong>91.79%</strong></td><td style="text-align:center"><strong>24,248 sent/s</strong></td></tr><tr><td style="text-align:left">🥉 + TurboQuant 2-bit</td><td style="text-align:center"><strong>2.50 MB</strong></td><td style="text-align:center"><strong>880x smaller</strong></td><td style="text-align:center"><strong>92.19%</strong></td><td style="text-align:center"><strong>20,013 sent/s</strong></td></tr></tbody></table>
<p>As the benchmarks indicate, the 2-bit model maintains performance levels comparable to the FP16 version, showing no significant loss in accuracy.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="step-1-distillation--teaching-a-lookup-table-to-understand-turkish">Step 1: Distillation — Teaching a Lookup Table to Understand Turkish<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#step-1-distillation--teaching-a-lookup-table-to-understand-turkish" class="hash-link" aria-label="Step 1: Distillation — Teaching a Lookup Table to Understand Turkish doğrudan bağlantı" title="Step 1: Distillation — Teaching a Lookup Table to Understand Turkish doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-architecture">The Architecture<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#the-architecture" class="hash-link" aria-label="The Architecture doğrudan bağlantı" title="The Architecture doğrudan bağlantı" translate="no">​</a></h3>
<p><a href="https://github.com/MinishLab/model2vec" target="_blank" rel="noopener noreferrer" class="">Model2Vec</a> is a technique that converts transformer sentence encoders into static embedding lookup tables. Instead of running 24 transformer layers per inference, you look up token embeddings from a matrix, average them, normalize, and you're done. O(1) per token, no attention heads, no GPU.</p>
<p>Our distillation pipeline works as follows:</p>
<p><strong>Teacher:</strong> We used the pre-computed 1024-dimensional BGE-M3 sentence embeddings for all 533,189 Turkish Wikipedia articles from the dataset: <a href="https://huggingface.co/datasets/hcsolakoglu/turkish_wikipedia_with_bge_m3_embeddings" target="_blank" rel="noopener noreferrer" class="">hcsolakoglu/turkish_wikipedia_with_bge_m3_embeddings</a>.</p>
<p><strong>Dimensionality Reduction:</strong> We fit a 256-dimensional PCA on the teacher embeddings, preserving 86.77% of the total variance. This gives us compact, decorrelated teacher targets.</p>
<p><strong>Student:</strong> A simple <code>nn.Embedding(39655, 256)</code> lookup table — just a matrix of shape 39,655 × 256, where 39,655 is the number of Turkish-relevant tokens we kept after pruning BGE-M3's multilingual vocabulary.</p>
<p><strong>Training Objective:</strong> For each Wikipedia article, tokenize with the pruned vocabulary, compute weighted mean pooling over student embeddings, L2-normalize, and minimize cosine distance to the PCA-projected teacher representation.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">Input Text → Tokenize → Lookup Embeddings → Weighted Mean Pool → L2 Normalize → Compare to Teacher</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-pruning-recovery-story">The Pruning Recovery Story<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#the-pruning-recovery-story" class="hash-link" aria-label="The Pruning Recovery Story doğrudan bağlantı" title="The Pruning Recovery Story doğrudan bağlantı" translate="no">​</a></h3>
<p>This was the hardest part. When you prune a 250k multilingual vocabulary down to ~39k Turkish tokens, many subword units disappear. The tokenizer starts segmenting words differently, and previously one-token words become multi-token sequences. We call this <strong>subword segmentation drift</strong>.</p>
<p>The unpruned Model2Vec model (using the full 373k vocabulary with a morphological expander called Akana) scored <strong>86.83%</strong> on STSb-TR at <strong>182 MB</strong>. After pruning to 39k tokens, the naïve warm-start model crashed to <strong>75.30%</strong> — an 11.5 point drop.</p>
<p>The key insight: <strong>don't freeze the embeddings; retrain everything from scratch.</strong> By training all 39,655 token embeddings end-to-end on 533k Wikipedia articles with gradient descent, the student model learns to compensate for segmentation drift. Each subword learns a representation that, when averaged with its new neighbors, produces the correct sentence-level embedding.</p>
<p>After 3 epochs of training, the pruned student model reached <strong>91.36%</strong> — beating the 182 MB morphology-expanded model by <strong>+4.53 points</strong> at 9.4x smaller size.</p>
<p><strong>Lesson learned:</strong> Explicit morphological expansion is unnecessary for Turkish Model2Vec. End-to-end distillation on real text heals the segmentation naturally.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="step-2-turboquant--why-crushing-precision-improves-quality">Step 2: TurboQuant — Why Crushing Precision Improves Quality<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#step-2-turboquant--why-crushing-precision-improves-quality" class="hash-link" aria-label="Step 2: TurboQuant — Why Crushing Precision Improves Quality doğrudan bağlantı" title="Step 2: TurboQuant — Why Crushing Precision Improves Quality doğrudan bağlantı" translate="no">​</a></h2>
<p>This is the surprising part. <a href="https://arxiv.org/abs/2501.11159" target="_blank" rel="noopener noreferrer" class="">Google's TurboQuant</a> (ICLR 2026) is a 3-stage vector quantization algorithm designed for KV-cache compression in LLMs. We adapted it for static sentence embeddings and discovered an unexpected benefit.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-turboquant-works">How TurboQuant Works<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#how-turboquant-works" class="hash-link" aria-label="How TurboQuant Works doğrudan bağlantı" title="How TurboQuant Works doğrudan bağlantı" translate="no">​</a></h3>
<p><strong>Stage 1 — Random Orthogonal Rotation:</strong></p>
<p>Before quantizing, TurboQuant multiplies the entire embedding matrix by a random orthogonal matrix <em>R</em> sampled from the Haar measure on O(d). This sounds like it should destroy information, but it actually <em>improves</em> the quantization properties.</p>
<p>Why? Embedding dimensions are typically anisotropic — a few dimensions carry most of the variance while most dimensions contain noise. The rotation redistributes variance equally across all dimensions, making each coordinate approximately Gaussian. This is critical for the next stage.</p>
<p><strong>Stage 2 — Lloyd-Max Scalar Quantization:</strong></p>
<p>Each rotated coordinate is independently quantized using optimal Lloyd-Max codebooks for the standard Gaussian distribution. For 4-bit quantization, this means 16 centroids per dimension; for 2-bit, just 4 centroids.</p>
<p>The Lloyd-Max quantizer minimizes mean squared error for Gaussian inputs. Because Stage 1 made all coordinates Gaussian, this is mathematically optimal — you get the minimum possible distortion for a given bit budget.</p>
<p><strong>Stage 3 (Disabled) — QJL Residual Correction:</strong></p>
<p>Google's original paper includes a 1-bit Quantized Johnson-Lindenstrauss sketch for the quantization residual. We found this <strong>hurts</strong> accuracy for sentence embeddings (1–2.5 point drop in Pearson <em>r</em>) because the JL error bound scales as O(1/√m), and at d=256, the sketch introduces more noise than it corrects.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-denoising-effect-2-bit-quantization-performance">The Denoising Effect: 2-bit Quantization Performance<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#the-denoising-effect-2-bit-quantization-performance" class="hash-link" aria-label="The Denoising Effect: 2-bit Quantization Performance doğrudan bağlantı" title="The Denoising Effect: 2-bit Quantization Performance doğrudan bağlantı" translate="no">​</a></h3>
<p>Here's the punchline. The Lloyd-Max quantizer helps maintain performance even at extremely low bitwidths:</p>
<ul>
<li class="">Small stochastic fluctuations in embedding coordinates (from gradient noise during training) sit near zero.</li>
<li class="">The coarse 2-bit quantizer maps all near-zero values to the same centroid, effectively snapping noise to zero.</li>
<li class="">Large, semantically meaningful values survive quantization because they fall into well-separated centroid bins.</li>
</ul>
<p>The result: quantization preserves the angular structure (cosine similarity) that matters for semantic tasks while significantly reducing the model size. This robust architecture ensures that our 2-bit model (<strong>92.19%</strong>) performs on par with the unquantized FP16 model (<strong>91.36%</strong>), demonstrating that high-quality embeddings can be maintained even under extreme compression.<br>
These results remain consistent across our evaluations on both the STSb-gorkem and STSb-emrecan datasets, demonstrating the robustness of our distilled model even under heavy compression.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="step-3-what-didnt-work-and-what-we-learned">Step 3: What Didn't Work (and What We Learned)<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#step-3-what-didnt-work-and-what-we-learned" class="hash-link" aria-label="Step 3: What Didn't Work (and What We Learned) doğrudan bağlantı" title="Step 3: What Didn't Work (and What We Learned) doğrudan bağlantı" translate="no">​</a></h2>
<p>Not everything was successful. Here's what we tried and discarded:</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="-principal-component-removal-sif">❌ Principal Component Removal (SIF)<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#-principal-component-removal-sif" class="hash-link" aria-label="❌ Principal Component Removal (SIF) doğrudan bağlantı" title="❌ Principal Component Removal (SIF) doğrudan bağlantı" translate="no">​</a></h3>
<p>The classic unsupervised trick of subtracting the first principal component. This hurts our model because PCA teacher targets are already centered and orthogonalized — PC removal removes the principal <em>semantic</em> direction learned from BGE-M3.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="-orthogonal-procrustes-alignment">❌ Orthogonal Procrustes Alignment<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#-orthogonal-procrustes-alignment" class="hash-link" aria-label="❌ Orthogonal Procrustes Alignment doğrudan bağlantı" title="❌ Orthogonal Procrustes Alignment doğrudan bağlantı" translate="no">​</a></h3>
<p>Post-hoc rotation alignment between student and teacher spaces. Not needed because end-to-end training absorbs the optimal rotation directly into the embedding weights.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="-qjl-residual-sketches">❌ QJL Residual Sketches<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#-qjl-residual-sketches" class="hash-link" aria-label="❌ QJL Residual Sketches doğrudan bağlantı" title="❌ QJL Residual Sketches doğrudan bağlantı" translate="no">​</a></h3>
<p>As described above. Designed for MIPS in ultra-high dimensions (d=4096+), not for cosine similarity in d=256.</p>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="try-it-yourself">Try It Yourself<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#try-it-yourself" class="hash-link" aria-label="Try It Yourself doğrudan bağlantı" title="Try It Yourself doğrudan bağlantı" translate="no">​</a></h2>
<p>All models are on Hugging Face and can be loaded in 2 lines of Python:</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="load-the-fp16-champion-1936-mb">Load the FP16 Champion (19.36 MB)<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#load-the-fp16-champion-1936-mb" class="hash-link" aria-label="Load the FP16 Champion (19.36 MB) doğrudan bağlantı" title="Load the FP16 Champion (19.36 MB) doğrudan bağlantı" translate="no">​</a></h3>
<div class="language-py codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-py codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> model2vec </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> StaticModel</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">model </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> StaticModel</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">from_pretrained</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"altaidevorg/turkish-bge-m3-model2vec"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">embeddings </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> model</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">encode</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token punctuation" style="color:rgb(212, 212, 212)">[</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token string" style="color:rgb(206, 145, 120)">"Yapay zeka modelleri doğal dil işlemede çığır açıyor."</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token string" style="color:rgb(206, 145, 120)">"İstanbul, Türkiye'nin en kalabalık şehridir."</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">]</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">print</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">embeddings</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">shape</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain">  </span><span class="token comment" style="color:rgb(106, 153, 85)"># (2, 256)</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="load-the-2-bit-turboquant-model-250-mb">Load the 2-bit TurboQuant Model (2.50 MB)<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#load-the-2-bit-turboquant-model-250-mb" class="hash-link" aria-label="Load the 2-bit TurboQuant Model (2.50 MB) doğrudan bağlantı" title="Load the 2-bit TurboQuant Model (2.50 MB) doğrudan bağlantı" translate="no">​</a></h3>
<div class="language-py codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-py codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> src</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">turboquant </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> TurboQuantStaticModel</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">tq </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> TurboQuantStaticModel</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">from_pretrained</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token string" style="color:rgb(206, 145, 120)">"altaidevorg/turkish-bge-m3-model2vec-turboquant-2bit"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">embeddings </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> tq</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">encode</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token punctuation" style="color:rgb(212, 212, 212)">[</span><span class="token string" style="color:rgb(206, 145, 120)">"Hafif modeller mobil cihazlarda harika çalışır."</span><span class="token punctuation" style="color:rgb(212, 212, 212)">]</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">print</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">embeddings</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">shape</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain">  </span><span class="token comment" style="color:rgb(106, 153, 85)"># (1, 256)</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="use-2-bit-turboquant-model-with-akana-250-mb">Use 2-bit TurboQuant Model with <a href="https://github.com/altaidevorg/akana" target="_blank" rel="noopener noreferrer" class="">Akana</a> (2.50 MB)<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#use-2-bit-turboquant-model-with-akana-250-mb" class="hash-link" aria-label="use-2-bit-turboquant-model-with-akana-250-mb doğrudan bağlantı" title="use-2-bit-turboquant-model-with-akana-250-mb doğrudan bağlantı" translate="no">​</a></h3>
<div class="language-py codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-py codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> akana</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token comment" style="color:rgb(106, 153, 85)"># TurboQuant 2-Bit Turkish Sentence Embeddings &amp; Semantic Similarity</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">vec </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> akana</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">embed</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"Türkiye'nin başkenti Ankara'dır."</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">print</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)">f"Vector dim: </span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token string-interpolation interpolation builtin" style="color:rgb(86, 156, 214)">len</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string-interpolation interpolation">vec</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)">"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain">  </span><span class="token comment" style="color:rgb(106, 153, 85)"># -&gt; 256</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token comment" style="color:rgb(106, 153, 85)"># Semantic cosine similarity</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">score </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> akana</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">similarity</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"ev"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token string" style="color:rgb(206, 145, 120)">"evler"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">print</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)">f"Similarity: </span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token string-interpolation interpolation">score</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token string-interpolation interpolation format-spec">.4f</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)">"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain">  </span><span class="token comment" style="color:rgb(106, 153, 85)"># -&gt; ~0.9130</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token comment" style="color:rgb(106, 153, 85)"># Batch embedding</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">vecs </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> akana</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">embed_batch</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token punctuation" style="color:rgb(212, 212, 212)">[</span><span class="token string" style="color:rgb(206, 145, 120)">"Merhaba dünya"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token string" style="color:rgb(206, 145, 120)">"Hava bugün çok güzel"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">]</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">print</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)">f"Batch size: </span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token string-interpolation interpolation builtin" style="color:rgb(86, 156, 214)">len</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string-interpolation interpolation">vecs</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)">"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain">  </span><span class="token comment" style="color:rgb(106, 153, 85)"># -&gt; 2</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="use-akana-from-the-cli">Use Akana from the CLI<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#use-akana-from-the-cli" class="hash-link" aria-label="Use Akana from the CLI doğrudan bağlantı" title="Use Akana from the CLI doğrudan bağlantı" translate="no">​</a></h3>
<div class="language-shell codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-shell codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token comment" style="color:rgb(106, 153, 85)"># Directly from cli</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">akana embed </span><span class="token string" style="color:rgb(206, 145, 120)">"Türkiye'nin başkenti Ankara'dır."</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">akana similarity </span><span class="token string" style="color:rgb(206, 145, 120)">"ev"</span><span class="token plain"> </span><span class="token string" style="color:rgb(206, 145, 120)">"evler"</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="load-the-4-bit-turboquant-model-492-mb">Load the 4-bit TurboQuant Model (4.92 MB)<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#load-the-4-bit-turboquant-model-492-mb" class="hash-link" aria-label="Load the 4-bit TurboQuant Model (4.92 MB) doğrudan bağlantı" title="Load the 4-bit TurboQuant Model (4.92 MB) doğrudan bağlantı" translate="no">​</a></h3>
<div class="language-py codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-py codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">tq4 </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> TurboQuantStaticModel</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">from_pretrained</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token string" style="color:rgb(206, 145, 120)">"altaidevorg/turkish-bge-m3-model2vec-turboquant-4bit"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="full-benchmark-results">Full Benchmark Results<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#full-benchmark-results" class="hash-link" aria-label="Full Benchmark Results doğrudan bağlantı" title="Full Benchmark Results doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="high-precision-models">High-Precision Models<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#high-precision-models" class="hash-link" aria-label="High-Precision Models doğrudan bağlantı" title="High-Precision Models doğrudan bağlantı" translate="no">​</a></h3>



























































<table><thead><tr><th style="text-align:left">Model</th><th style="text-align:center">Vocab</th><th style="text-align:center">Size</th><th style="text-align:center">STSb (gorkem)</th><th style="text-align:center">STSb (emrecan)</th><th style="text-align:center">Speed</th><th style="text-align:center">Speedup</th></tr></thead><tbody><tr><td style="text-align:left">BGE-M3 (Teacher)</td><td style="text-align:center">250,002</td><td style="text-align:center">~2,200 MB</td><td style="text-align:center">96.35%</td><td style="text-align:center">79.57%</td><td style="text-align:center">79 s/s</td><td style="text-align:center">1.0x</td></tr><tr><td style="text-align:left"><strong>533k Distilled Champion</strong></td><td style="text-align:center"><strong>39,655</strong></td><td style="text-align:center"><strong>19.36 MB</strong></td><td style="text-align:center"><strong>91.36%</strong></td><td style="text-align:center"><strong>64.34%</strong></td><td style="text-align:center"><strong>63,012 s/s</strong></td><td style="text-align:center"><strong>797x</strong></td></tr><tr><td style="text-align:left">150k POC Distilled</td><td style="text-align:center">32,873</td><td style="text-align:center">16.05 MB</td><td style="text-align:center">82.83%</td><td style="text-align:center">58.82%</td><td style="text-align:center">62,403 s/s</td><td style="text-align:center">789x</td></tr><tr><td style="text-align:left">Original Model2Vec + Akana</td><td style="text-align:center">373,489</td><td style="text-align:center">182.37 MB</td><td style="text-align:center">86.83%</td><td style="text-align:center">69.70%</td><td style="text-align:center">61,284 s/s</td><td style="text-align:center">775x</td></tr><tr><td style="text-align:left">Pruned Baseline (Untrained)</td><td style="text-align:center">32,873</td><td style="text-align:center">16.05 MB</td><td style="text-align:center">75.30%</td><td style="text-align:center">54.28%</td><td style="text-align:center">52,459 s/s</td><td style="text-align:center">664x</td></tr></tbody></table>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="edge--on-device-models-turboquant">Edge &amp; On-Device Models (TurboQuant)<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#edge--on-device-models-turboquant" class="hash-link" aria-label="Edge &amp; On-Device Models (TurboQuant) doğrudan bağlantı" title="Edge &amp; On-Device Models (TurboQuant) doğrudan bağlantı" translate="no">​</a></h3>


















































<table><thead><tr><th style="text-align:left">Model</th><th style="text-align:center">Bits</th><th style="text-align:center">Size</th><th style="text-align:center">Compression</th><th style="text-align:center">STSb (gorkem)</th><th style="text-align:center">STSb (emrecan)</th><th style="text-align:center">Speed</th></tr></thead><tbody><tr><td style="text-align:left"><strong>Champion + TQ 4-bit</strong></td><td style="text-align:center">INT4</td><td style="text-align:center"><strong>4.92 MB</strong></td><td style="text-align:center"><strong>447x</strong></td><td style="text-align:center"><strong>91.79%</strong></td><td style="text-align:center"><strong>64.19%</strong></td><td style="text-align:center"><strong>24,248 s/s</strong></td></tr><tr><td style="text-align:left"><strong>Champion + TQ 2-bit</strong></td><td style="text-align:center">INT2</td><td style="text-align:center"><strong>2.50 MB</strong></td><td style="text-align:center"><strong>880x</strong></td><td style="text-align:center"><strong>92.19%</strong></td><td style="text-align:center"><strong>63.53%</strong></td><td style="text-align:center"><strong>20,013 s/s</strong></td></tr><tr><td style="text-align:left">150k POC + TQ 4-bit</td><td style="text-align:center">INT4</td><td style="text-align:center">4.08 MB</td><td style="text-align:center">539x</td><td style="text-align:center">81.59%</td><td style="text-align:center">56.89%</td><td style="text-align:center">24,549 s/s</td></tr><tr><td style="text-align:left">150k POC + TQ 2-bit</td><td style="text-align:center">INT2</td><td style="text-align:center">2.07 MB</td><td style="text-align:center">1,062x</td><td style="text-align:center">80.61%</td><td style="text-align:center">56.29%</td><td style="text-align:center">23,121 s/s</td></tr></tbody></table>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="conclusion">Conclusion<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#conclusion" class="hash-link" aria-label="Conclusion doğrudan bağlantı" title="Conclusion doğrudan bağlantı" translate="no">​</a></h2>
<p>We compressed a 2.2 GB multilingual transformer into a 2.5 MB static embedding model that:</p>
<ul>
<li class="">Runs <strong>800x faster</strong> on CPU</li>
<li class="">Retains <strong>92% of teacher quality</strong> on Turkish STS</li>
<li class="">Requires <strong>zero GPU, zero morphological tools, zero external dependencies</strong></li>
<li class="">Fits inside a <strong>mobile app, a browser extension, or a smartwatch</strong></li>
</ul>
<p>The two key insights:</p>
<ol>
<li class=""><strong>End-to-end distillation on real text beats morphological engineering.</strong> Don't enumerate surface forms; let gradient descent learn what subwords mean in context.</li>
<li class=""><strong>Quantization can be a regularizer, not just compression.</strong> TurboQuant's random rotation + coarse quantization acts as an implicit denoiser that removes training noise while preserving semantic structure.</li>
</ol>
<hr>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="links">Links<a href="https://altai.dev/tr/blog/turkish-sentence-embeddings-model2vec-turboquant#links" class="hash-link" aria-label="Links doğrudan bağlantı" title="Links doğrudan bağlantı" translate="no">​</a></h2>
<ul>
<li class=""><strong>GitHub:</strong> <a href="https://github.com/altaidevorg/model2vec_experiments" target="_blank" rel="noopener noreferrer" class="">github.com/altaidevorg/model2vec_experiments</a></li>
<li class=""><strong>🥇 FP16 Model (19.36 MB):</strong> <a href="https://huggingface.co/altaidevorg/turkish-bge-m3-model2vec" target="_blank" rel="noopener noreferrer" class="">altaidevorg/turkish-bge-m3-model2vec</a></li>
<li class=""><strong>🥈 4-bit Model (4.92 MB):</strong> <a href="https://huggingface.co/altaidevorg/turkish-bge-m3-model2vec-turboquant-4bit" target="_blank" rel="noopener noreferrer" class="">altaidevorg/turkish-bge-m3-model2vec-turboquant-4bit</a></li>
<li class=""><strong>🥉 2-bit Model (2.50 MB):</strong> <a href="https://huggingface.co/altaidevorg/turkish-bge-m3-model2vec-turboquant-2bit" target="_blank" rel="noopener noreferrer" class="">altaidevorg/turkish-bge-m3-model2vec-turboquant-2bit</a></li>
<li class=""><strong>Training Dataset:</strong> <a href="https://huggingface.co/datasets/hcsolakoglu/turkish_wikipedia_with_bge_m3_embeddings" target="_blank" rel="noopener noreferrer" class="">hcsolakoglu/turkish_wikipedia_with_bge_m3_embeddings</a></li>
</ul>
<hr>
<p><em>Built by <a href="https://github.com/altaidevorg" target="_blank" rel="noopener noreferrer" class="">Altai</a>. Licensed under Apache 2.0.</em></p>]]></content>
        <author>
            <name>Hakan Doğan</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[OpenSimula examples walkthrough]]></title>
        <id>https://altai.dev/tr/blog/simula-example</id>
        <link href="https://altai.dev/tr/blog/simula-example"/>
        <updated>2026-04-29T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[A guided explanation of the OpenSimula pieces used in examples/simula.]]></summary>
        <content type="html"><![CDATA[<p>This walkthrough explains the scripts in <code>examples/simula/</code> as OpenSimula
compositions, not just as commands to run. The goal is to make the moving parts
visible so you can replace the policy examples with your own documents, task
format, checkpoint layout, model provider, or downstream export path.</p>
<p>The examples exercise <code>afterimage.simula</code>, an experimental open implementation
of Simula-style synthetic data mechanisms inspired by Davidson et al.,
<em>Reasoning-Driven Synthetic Data Generation and Evaluation</em>. They are not a
Google reference implementation. The examples default to <code>gemini-2.5-flash</code>
because it is fast enough for iterative taxonomy, scenario, and critic loops
while staying close to the teacher-model family used in the paper.</p>
<p>There are three runnable scripts:</p>





















<table><thead><tr><th>Script</th><th>What it demonstrates</th></tr></thead><tbody><tr><td><code>minimal_pipeline.py</code></td><td>One document-grounded single-QA datapoint with taxonomy, strategy sampling, meta-prompting, and requirement-critic refinement.</td></tr><tr><td><code>mcq_pipeline.py</code></td><td>One four-option MCQ datapoint with the same global/local pipeline plus the MCQ double-critic gate.</td></tr><tr><td><code>corpus_batch_qa.py</code></td><td>A production-shaped batch run with a larger static corpus, checkpoint files, resume support, bounded concurrency, JSONL append, and optional Hub upload.</td></tr></tbody></table>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-mental-model">The Mental Model<a href="https://altai.dev/tr/blog/simula-example#the-mental-model" class="hash-link" aria-label="The Mental Model doğrudan bağlantı" title="The Mental Model doğrudan bağlantı" translate="no">​</a></h2>
<p>OpenSimula separates global coverage from local variety:</p>













































<table><thead><tr><th>Phase</th><th>OpenSimula piece</th><th>Purpose</th></tr></thead><tbody><tr><td>Dataset intent</td><td><code>instruction_y</code></td><td>Describes the target domain, audience, format, constraints, and non-goals.</td></tr><tr><td>Optional source grounding</td><td><code>DocumentProvider</code></td><td>Supplies bounded excerpts used while creating factors and taxonomies.</td></tr><tr><td>Global diversification</td><td><code>OpenSimula.build_taxonomy()</code></td><td>Builds factor taxonomies: the conceptual space the dataset should cover.</td></tr><tr><td>Joint sampling</td><td><code>OpenSimula.infer_strategies()</code> and <code>sample_mix()</code></td><td>Chooses compatible taxonomy nodes to combine into one datapoint requirement mix.</td></tr><tr><td>Local diversification</td><td><code>OpenSimula.draw_meta_prompt()</code></td><td>Generates scenario/meta-prompt candidates for a sampled mix, then optionally complexifies one.</td></tr><tr><td>Task generation</td><td><code>generate_single_qa_datapoint()</code> or <code>generate_mcq_datapoint()</code></td><td>Produces a row and checks it against the requirements.</td></tr><tr><td>Persistence</td><td><code>Checkpointer</code> and <code>append_datapoints_jsonl()</code></td><td>Saves reusable taxonomy/strategy artifacts and accepted datapoints.</td></tr></tbody></table>
<p>The generation loop in the minimal single-QA path looks like this:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">instruction_y + optional documents</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; build_taxonomy(y, S, D, N)</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; infer_strategies(bundle)</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; sample_mix(bundle, strategy)</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; draw_meta_prompt(y, bundle, mix, K, complexify_c)</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; generate_single_qa_datapoint(...)</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; requirement critic accepts or refines</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; DataPointRecord JSON</span><br></div></code></pre></div></div>
<p>For MCQ, the final task step adds a second gate:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">MCQ JSON</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; requirement critic/refine loop</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; double-critic probes the labeled answer</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; accepted DataPointRecord or None</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="imports-the-building-blocks">Imports: The Building Blocks<a href="https://altai.dev/tr/blog/simula-example#imports-the-building-blocks" class="hash-link" aria-label="Imports: The Building Blocks doğrudan bağlantı" title="Imports: The Building Blocks doğrudan bağlantı" translate="no">​</a></h2>
<p>The examples all start by creating an LLM provider:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> afterimage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">providers </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> LLMFactory</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">llm </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> LLMFactory</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">create</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    provider</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token string" style="color:rgb(206, 145, 120)">"gemini"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    model_name</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token string" style="color:rgb(206, 145, 120)">"gemini-2.5-flash"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    api_key</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">api_key</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p><code>OpenSimula</code> takes that provider and uses it for every structured LLM step:
taxonomy proposals, critic merges, strategy inference, meta-prompts, task JSON,
requirement critiques, refinements, and MCQ double-critic probes.</p>
<p>The smallest examples import:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> afterimage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">providers </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> InMemoryDocumentProvider</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> LLMFactory</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> afterimage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">simula </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> OpenSimula</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> configure_example_console</span><br></div></code></pre></div></div>
<p>You can read them as roles:</p>
<ul>
<li class=""><code>LLMFactory</code> creates the provider used by the OpenSimula facade.</li>
<li class=""><code>InMemoryDocumentProvider</code> supplies the example policy excerpts as source
material.</li>
<li class=""><code>OpenSimula</code> orchestrates taxonomy, sampling, meta-prompt, and datapoint
generation.</li>
<li class=""><code>configure_example_console()</code> keeps progress output readable by muting noisy
third-party INFO logs.</li>
</ul>
<p>The batch script adds persistence helpers:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> afterimage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">simula </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    Checkpointer</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    OpenSimulaRunConfig</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    append_datapoints_jsonl</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    load_checkpoint</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>These helpers are the difference between a toy single-row run and a reusable
workflow. They let you save the expensive global scaffold once, resume from it,
and append accepted task rows as generation completes.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="dataset-intent-instruction_y">Dataset Intent: <code>instruction_y</code><a href="https://altai.dev/tr/blog/simula-example#dataset-intent-instruction_y" class="hash-link" aria-label="dataset-intent-instruction_y doğrudan bağlantı" title="dataset-intent-instruction_y doğrudan bağlantı" translate="no">​</a></h2>
<p>Each script defines one <code>INSTRUCTION_Y</code> string. This is the global dataset
specification, usually called <code>y</code> in the paper mapping.</p>
<p>In <code>minimal_pipeline.py</code>, the intent is:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">INSTRUCTION_Y </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token triple-quoted-string string" style="color:rgb(206, 145, 120)">"""\</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token triple-quoted-string string" style="color:rgb(206, 145, 120)">You are generating synthetic **training Q&amp;A** for enterprise employees...</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token triple-quoted-string string" style="color:rgb(206, 145, 120)">"""</span><br></div></code></pre></div></div>
<p>The wording does a lot of work:</p>
<ul>
<li class="">It names the audience: enterprise employees.</li>
<li class="">It names the task format: training Q&amp;A.</li>
<li class="">It establishes grounding: answers must use the policy excerpts.</li>
<li class="">It sets style and size constraints: factual tone, no panic language, word
limits.</li>
<li class="">It prevents unwanted hallucination: no vendor-specific products or laws that
are not mentioned.</li>
</ul>
<p>When adapting these examples, this is usually the first thing to rewrite. A good
<code>instruction_y</code> is specific enough that an evaluator could reject rows that miss
the mark.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="optional-documents-source-grounding">Optional Documents: Source Grounding<a href="https://altai.dev/tr/blog/simula-example#optional-documents-source-grounding" class="hash-link" aria-label="Optional Documents: Source Grounding doğrudan bağlantı" title="Optional Documents: Source Grounding doğrudan bağlantı" translate="no">​</a></h2>
<p><code>minimal_pipeline.py</code> and <code>corpus_batch_qa.py</code> use static policy excerpts:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">docs </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> InMemoryDocumentProvider</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">POLICY_EXCERPTS</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>Those excerpts are used while building the taxonomy. They help OpenSimula infer
factors and branches that are actually present in the domain material. In the
minimal script, the source material is intentionally tiny. In the batch script,
<code>CORPUS_EXCERPTS</code> has six policy-style snippets so the taxonomy has more
surface area.</p>
<p><code>mcq_pipeline.py</code> passes:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">document_provider</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token boolean">None</span><br></div></code></pre></div></div>
<p>That script demonstrates the no-corpus path. The taxonomy is inferred from
<code>instruction_y</code> alone, which is useful when the target domain is a format or
assessment style rather than a closed set of documents.</p>
<p>For real runs, replace <code>InMemoryDocumentProvider</code> with a provider that matches
your source material. Directory, JSONL, or vector-backed providers are better
fits once the corpus no longer fits comfortably in a Python list.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="console-setup-and-progress">Console Setup and Progress<a href="https://altai.dev/tr/blog/simula-example#console-setup-and-progress" class="hash-link" aria-label="Console Setup and Progress doğrudan bağlantı" title="Console Setup and Progress doğrudan bağlantı" translate="no">​</a></h2>
<p>The examples call:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">configure_example_console</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>OpenSimula makes many structured LLM calls. Without quiet logging, HTTP client
and SDK messages can bury the useful progress signal. The examples also pass:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">show_progress</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token boolean">True</span><br></div></code></pre></div></div>
<p>to <code>build_taxonomy()</code>. That enables nested <code>tqdm</code> progress for factor proposal
and per-factor breadth-first taxonomy expansion.</p>
<p>If a script appears paused after constructing <code>OpenSimula</code>, it is usually inside
<code>build_taxonomy()</code>. The first structured call can take tens of seconds, and
taxonomy expansion is sequential by design because each critic and planning step
depends on the current tree state.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="taxonomy-construction">Taxonomy Construction<a href="https://altai.dev/tr/blog/simula-example#taxonomy-construction" class="hash-link" aria-label="Taxonomy Construction doğrudan bağlantı" title="Taxonomy Construction doğrudan bağlantı" translate="no">​</a></h2>
<p>All three scripts build a <code>TaxonomyBundle</code>:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">bundle </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> sim</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">build_taxonomy</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    INSTRUCTION_Y</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    document_provider</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">docs</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    target_depth_D</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">TARGET_DEPTH_D</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    proposal_N</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">PROPOSAL_N</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    max_factors</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">MAX_FACTORS</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    max_children_per_node</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">MAX_CHILDREN_PER_NODE</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    max_frontier_per_depth</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">MAX_FRONTIER_PER_DEPTH</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    show_progress</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token boolean">True</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">OpenSimula</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">validate_taxonomy_bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>Important knobs:</p>

































<table><thead><tr><th>Argument</th><th>Meaning</th></tr></thead><tbody><tr><td><code>target_depth_D</code></td><td>Maximum taxonomy depth per factor. Deeper trees create finer control but cost more.</td></tr><tr><td><code>proposal_N</code></td><td>Number of independent child proposals before a critic merges them.</td></tr><tr><td><code>max_factors</code></td><td>Caps how many top-level factors are expanded.</td></tr><tr><td><code>max_children_per_node</code></td><td>Caps breadth after the critic merge step.</td></tr><tr><td><code>max_frontier_per_depth</code></td><td>Caps how many nodes are expanded at each depth.</td></tr><tr><td><code>show_progress</code></td><td>Shows <code>tqdm</code> progress during the expensive taxonomy phase.</td></tr></tbody></table>
<p>The caps matter. A taxonomy is a branching structure; allowing too many factors,
children, or frontier nodes can multiply into hundreds of sequential LLM calls.
The examples choose conservative defaults so local runs are predictable.</p>
<p>The resulting bundle stores:</p>
<ul>
<li class="">the original <code>instruction_y</code>;</li>
<li class="">document digests for bounded excerpts used during construction;</li>
<li class="">accepted factors;</li>
<li class="">one taxonomy tree per factor;</li>
<li class="">expansion traces and per-depth plans for auditability.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="sampling-strategies-and-mixes">Sampling Strategies and Mixes<a href="https://altai.dev/tr/blog/simula-example#sampling-strategies-and-mixes" class="hash-link" aria-label="Sampling Strategies and Mixes doğrudan bağlantı" title="Sampling Strategies and Mixes doğrudan bağlantı" translate="no">​</a></h2>
<p>After the taxonomy is built, the scripts infer sampling strategies:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">spec </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> sim</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">infer_strategies</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">mix </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> sim</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">sample_mix</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> spec</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>The strategy spec describes which factors can be sampled together and with what
weights. This prevents arbitrary combinations of taxonomy leaves from producing
incoherent requirements.</p>
<p><code>sample_mix()</code> draws one concrete <code>Mix</code>: a tuple of taxonomy nodes that becomes
the requirement set for the next datapoint. In batch mode, every sample gets its
own independent mix.</p>
<p>If you need full control, you can hand-author a <code>SamplingStrategySpec</code>, but the
examples use <code>infer_strategies()</code> because it is the fastest way to get a
reasonable mechanism from a new taxonomy.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="meta-prompts-local-diversity">Meta-Prompts: Local Diversity<a href="https://altai.dev/tr/blog/simula-example#meta-prompts-local-diversity" class="hash-link" aria-label="Meta-Prompts: Local Diversity doğrudan bağlantı" title="Meta-Prompts: Local Diversity doğrudan bağlantı" translate="no">​</a></h2>
<p>The sampled mix defines requirements. The meta-prompt turns those requirements
into a local scenario:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">meta </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> sim</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">draw_meta_prompt</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    instruction_y</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">instruction_y</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    bundle</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    mix</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">mix</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    K</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">META_PROMPT_K</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    complexify_c</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">COMPLEXIFY_C</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    sequential</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token boolean">False</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p><code>K</code> controls how many candidate scenarios are generated before one is
subsampled. Larger values give the model more chances to vary framing while
keeping the same global requirements.</p>
<p><code>complexify_c</code> is the probability of running an additional complexification
step. It is orthogonal to taxonomy coverage: complexity changes the difficulty
or nuance of a scenario, while the mix controls which conceptual requirements
are present.</p>
<p>The examples set <code>sequential=False</code> for speed. The sequential path generates
scenario candidates one by one with prior attempts in context, which can reduce
mode collapse at higher cost.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="single-qa-generation">Single-QA Generation<a href="https://altai.dev/tr/blog/simula-example#single-qa-generation" class="hash-link" aria-label="Single-QA Generation doğrudan bağlantı" title="Single-QA Generation doğrudan bağlantı" translate="no">​</a></h2>
<p><code>minimal_pipeline.py</code> uses:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">row </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> sim</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">generate_single_qa_datapoint</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    instruction_y</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">instruction_y</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    bundle</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    mix</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">mix</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    meta</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">meta</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>This produces one <code>DataPointRecord</code> with a single question and answer. The task
generator first emits task JSON, then the requirement critic checks whether the
row satisfies <code>instruction_y</code>, the sampled mix, and the meta-prompt.</p>
<p>If the critic rejects the row, OpenSimula can ask the LLM to refine it. The
default maximum is four refine rounds. If the row still does not satisfy the
requirements, the method returns <code>None</code>.</p>
<p>That is why the script handles both outcomes:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">if</span><span class="token plain"> row </span><span class="token keyword" style="color:rgb(86, 156, 214)">is</span><span class="token plain"> </span><span class="token boolean">None</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token keyword" style="color:rgb(86, 156, 214)">print</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"No row accepted..."</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">else</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token keyword" style="color:rgb(86, 156, 214)">print</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">row</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">model_dump_json</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">indent</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token number" style="color:rgb(181, 206, 168)">2</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>For a real dataset, treat <code>None</code> as a normal rejected sample, not an exception.
Batch pipelines should count accepted rows and continue.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="mcq-generation-and-double-critic">MCQ Generation and Double-Critic<a href="https://altai.dev/tr/blog/simula-example#mcq-generation-and-double-critic" class="hash-link" aria-label="MCQ Generation and Double-Critic doğrudan bağlantı" title="MCQ Generation and Double-Critic doğrudan bağlantı" translate="no">​</a></h2>
<p><code>mcq_pipeline.py</code> follows the same taxonomy, strategy, mix, and meta-prompt
path. The task-specific call changes:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">row </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> sim</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">generate_mcq_datapoint</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    instruction_y</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">instruction_y</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    bundle</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    mix</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">mix</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    meta</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">meta</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    num_choices</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">NUM_CHOICES</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>The MCQ path still runs the requirement critic and refinement loop. After that,
it adds a double-critic gate for labeled-answer quality. The gate asks two
independent structured probes about the answer label: one from the "this is
correct" angle and one from the "this is incorrect" angle. A row is accepted
only when the probes support a verifiable single correct answer.</p>
<p>Use this path when the output label matters as much as the question text. For
free-form QA, the single-QA requirement critic is usually enough.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="batch-generation">Batch Generation<a href="https://altai.dev/tr/blog/simula-example#batch-generation" class="hash-link" aria-label="Batch Generation doğrudan bağlantı" title="Batch Generation doğrudan bağlantı" translate="no">​</a></h2>
<p><code>corpus_batch_qa.py</code> turns the single-QA pattern into a multi-sample workflow:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">async</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(86, 156, 214)">for</span><span class="token plain"> _idx</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> rec </span><span class="token keyword" style="color:rgb(86, 156, 214)">in</span><span class="token plain"> sim</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">aiter_single_qa_samples</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    instruction_y</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">instruction_y</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    bundle</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    spec</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">spec</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    n</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">num_samples</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    K</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">META_PROMPT_K</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    complexify_c</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">COMPLEXIFY_C</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    sequential</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token boolean">False</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    max_concurrency</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">max_concurrency</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    rng</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">rng</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token keyword" style="color:rgb(86, 156, 214)">if</span><span class="token plain"> rec </span><span class="token keyword" style="color:rgb(86, 156, 214)">is</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(86, 156, 214)">not</span><span class="token plain"> </span><span class="token boolean">None</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">        append_datapoints_jsonl</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">jsonl_path</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">[</span><span class="token plain">rec</span><span class="token punctuation" style="color:rgb(212, 212, 212)">]</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>Each sample independently draws:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">mix -&gt; meta-prompt -&gt; task row -&gt; critic/refine</span><br></div></code></pre></div></div>
<p><code>max_concurrency</code> bounds how many sample pipelines run at once. Increasing it
can improve throughput, but each worker still performs several LLM calls. Start
with <code>1</code> or <code>2</code> when using tight provider quotas.</p>
<p>Rows are appended as each async sample completes. This is deliberate: if a long
run crashes after accepting some rows, those rows are already in
<code>data/train.jsonl</code>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="checkpoints-and-resume">Checkpoints and Resume<a href="https://altai.dev/tr/blog/simula-example#checkpoints-and-resume" class="hash-link" aria-label="Checkpoints and Resume doğrudan bağlantı" title="Checkpoints and Resume doğrudan bağlantı" translate="no">​</a></h2>
<p>The batch script saves the expensive global artifacts under <code>opensimula/</code>:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">&lt;output-dir&gt;/</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  opensimula/</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    manifest.json</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    taxonomy_bundle.json</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    sampling_strategy.json</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    run_config.json</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  data/</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    train.jsonl</span><br></div></code></pre></div></div>
<p>The save path is:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">with</span><span class="token plain"> Checkpointer</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">out</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(86, 156, 214)">as</span><span class="token plain"> cp</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    bundle</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">save</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">cp</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    spec</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">save</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">cp</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    cp</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">write_run_config</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">run_cfg</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p><code>run_config.json</code> is typed as <code>OpenSimulaRunConfig</code>; it records the model,
temperature, taxonomy knobs, sample count, seed, output path, and corpus size.
The <code>manifest.json</code> identifies the subtree as an AfterImage <code>opensimula</code>
checkpoint.</p>
<p>On resume, the script skips taxonomy construction and strategy inference when
possible:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">ckpt </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> load_checkpoint</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">out</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">bundle </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> ckpt</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">bundle</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">spec </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> ckpt</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">sampling_strategy</span><br></div></code></pre></div></div>
<p>Use <code>--resume</code> when you want to append more rows using the same conceptual
scaffold. This keeps follow-up generation cheaper and makes dataset expansion
more consistent.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="running-the-examples">Running the Examples<a href="https://altai.dev/tr/blog/simula-example#running-the-examples" class="hash-link" aria-label="Running the Examples doğrudan bağlantı" title="Running the Examples doğrudan bağlantı" translate="no">​</a></h2>
<p>All scripts require <code>GEMINI_API_KEY</code> unless you edit the provider setup:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token builtin class-name" style="color:rgb(78, 201, 176)">export</span><span class="token plain"> </span><span class="token assign-left variable" style="color:rgb(156, 220, 254)">GEMINI_API_KEY</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token string" style="color:rgb(206, 145, 120)">"your_api_key_here"</span><br></div></code></pre></div></div>
<p>From <code>examples/simula/</code>:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">python minimal_pipeline.py</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">python mcq_pipeline.py</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">python corpus_batch_qa.py --output-dir ./runs/corpus1 --num-samples </span><span class="token number" style="color:rgb(181, 206, 168)">8</span><span class="token plain"> --max-concurrency </span><span class="token number" style="color:rgb(181, 206, 168)">2</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">python corpus_batch_qa.py --output-dir ./runs/corpus1 </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">--resume</span><span class="token plain"> --num-samples </span><span class="token number" style="color:rgb(181, 206, 168)">4</span><br></div></code></pre></div></div>
<p>For Hub upload from the batch script, set <code>HF_TOKEN</code> or
<code>HUGGINGFACE_HUB_TOKEN</code>, then pass:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">python corpus_batch_qa.py --output-dir ./runs/corpus1 --push-hf username/dataset-name</span><br></div></code></pre></div></div>
<p>The push path uploads the <code>opensimula/</code> checkpoint subtree and a generated
dataset README. It does not replace inspecting your local JSONL before scaling.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-to-adapt-this-example">How To Adapt This Example<a href="https://altai.dev/tr/blog/simula-example#how-to-adapt-this-example" class="hash-link" aria-label="How To Adapt This Example doğrudan bağlantı" title="How To Adapt This Example doğrudan bağlantı" translate="no">​</a></h2>
<p>Use this checklist when adapting <code>examples/simula/</code> to another domain:</p>
<ol>
<li class="">Rewrite <code>INSTRUCTION_Y</code> so it states the audience, task format, constraints,
and non-goals.</li>
<li class="">Replace the static policy excerpts with realistic source material, or pass
<code>document_provider=None</code> for an instruction-only taxonomy.</li>
<li class="">Start with <code>target_depth_D=2</code>, <code>proposal_N=2</code> or <code>3</code>, and capped factors while
you inspect the taxonomy.</li>
<li class="">Validate the bundle after construction.</li>
<li class="">Infer strategies once per taxonomy bundle, then sample mixes many times.</li>
<li class="">Tune <code>META_PROMPT_K</code> and <code>COMPLEXIFY_C</code> only after the basic rows look
grounded and on-format.</li>
<li class="">Use single-QA for free-form answers; use MCQ when labeled choices need the
double-critic gate.</li>
<li class="">Save checkpoints for any run whose taxonomy cost you do not want to repeat.</li>
<li class="">Append accepted rows incrementally and treat rejected rows as expected
sampling loss.</li>
<li class="">Inspect <code>data/train.jsonl</code> before increasing <code>num_samples</code> or concurrency.</li>
</ol>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="common-variations">Common Variations<a href="https://altai.dev/tr/blog/simula-example#common-variations" class="hash-link" aria-label="Common Variations doğrudan bağlantı" title="Common Variations doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="use-files-instead-of-inline-strings">Use Files Instead of Inline Strings<a href="https://altai.dev/tr/blog/simula-example#use-files-instead-of-inline-strings" class="hash-link" aria-label="Use Files Instead of Inline Strings doğrudan bağlantı" title="Use Files Instead of Inline Strings doğrudan bağlantı" translate="no">​</a></h3>
<p>Replace <code>InMemoryDocumentProvider</code> with a provider that reads your corpus. Keep
the rest of the OpenSimula flow the same:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">documents -&gt; taxonomy -&gt; strategies -&gt; mixes -&gt; meta-prompts -&gt; datapoints</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="generate-a-larger-qa-dataset">Generate a Larger QA Dataset<a href="https://altai.dev/tr/blog/simula-example#generate-a-larger-qa-dataset" class="hash-link" aria-label="Generate a Larger QA Dataset doğrudan bağlantı" title="Generate a Larger QA Dataset doğrudan bağlantı" translate="no">​</a></h3>
<p>Use <code>corpus_batch_qa.py</code> as the starting point. Increase <code>--num-samples</code>, keep
<code>--max-concurrency</code> conservative, and use <code>--resume</code> so you do not rebuild the
taxonomy for every extension run.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="build-a-multiple-choice-benchmark">Build a Multiple-Choice Benchmark<a href="https://altai.dev/tr/blog/simula-example#build-a-multiple-choice-benchmark" class="hash-link" aria-label="Build a Multiple-Choice Benchmark doğrudan bağlantı" title="Build a Multiple-Choice Benchmark doğrudan bağlantı" translate="no">​</a></h3>
<p>Start from <code>mcq_pipeline.py</code>, strengthen <code>INSTRUCTION_Y</code> around answer
verifiability, and keep <code>num_choices=4</code> unless your downstream benchmark format
requires a different shape.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="feed-simula-scenarios-into-multi-turn-conversations">Feed Simula Scenarios Into Multi-Turn Conversations<a href="https://altai.dev/tr/blog/simula-example#feed-simula-scenarios-into-multi-turn-conversations" class="hash-link" aria-label="Feed Simula Scenarios Into Multi-Turn Conversations doğrudan bağlantı" title="Feed Simula Scenarios Into Multi-Turn Conversations doğrudan bağlantı" translate="no">​</a></h3>
<p>OpenSimula can also prepare first-turn scenarios for <code>ConversationGenerator</code> via
<code>SimulaInstructionGeneratorCallback</code>. That bridge does not call the LLM itself;
it replays precomputed scenario text and metadata so the normal AfterImage
conversation machinery can generate multi-turn assistant/user dialogs.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-these-examples-teach">What These Examples Teach<a href="https://altai.dev/tr/blog/simula-example#what-these-examples-teach" class="hash-link" aria-label="What These Examples Teach doğrudan bağlantı" title="What These Examples Teach doğrudan bağlantı" translate="no">​</a></h2>
<p>The reusable lesson is not "security policy Q&amp;A." The reusable pattern is:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">dataset intent</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; optional source grounding</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; factor taxonomies for global coverage</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; weighted strategy mixes</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; local meta-prompts for scenario diversity</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; task-specific generation and critics</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; checkpointed artifacts plus accepted JSONL rows</span><br></div></code></pre></div></div>
<p>Once you understand that shape, you can apply OpenSimula to employee training,
technical assessments, policy education, support simulations, benchmark-style
MCQ generation, or any domain where you want explicit control over both coverage
and local variation.</p>]]></content>
        <author>
            <name>Hakan Doğan</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Convert your raw corpus to SFT data: A walkthrough with AfterImage to generate legal research conversations]]></title>
        <id>https://altai.dev/tr/blog/caselaw-rag-example</id>
        <link href="https://altai.dev/tr/blog/caselaw-rag-example"/>
        <updated>2026-04-28T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[A guided explanation of the AfterImage pieces used in examples/caselaw_rag/generate.py.]]></summary>
        <content type="html"><![CDATA[<p>This walkthrough explains <code>examples/caselaw_rag/generate.py</code> as an AfterImage
composition, not just as a script to run. The goal is to make each moving part
visible so you can replace the caselaw pieces with your own documents, prompts,
storage, retrieval backend, or model provider.</p>
<p>The script generates synthetic legal research conversations. For demonstration,
we use a small slice of the
<a href="https://huggingface.co/datasets/free-law/Caselaw_Access_Project_embeddings" target="_blank" rel="noopener noreferrer" class=""><code>free-law/Caselaw_Access_Project_embeddings</code></a>
dataset: cleaned U.S. court opinion text with precomputed
<code>BAAI/bge-base-en-v1.5</code> vectors. This keeps the example easy to run while still
showing the full RAG data-generation pattern.</p>
<p>You can use the same structure with your own corpus. The important requirement
is that <code>index_corpus.py</code> or your own ingestion job creates a Qdrant collection
with document text in a known payload field and vectors that match the embedding
model used later for query-time retrieval.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-mental-model">The Mental Model<a href="https://altai.dev/tr/blog/caselaw-rag-example#the-mental-model" class="hash-link" aria-label="The Mental Model doğrudan bağlantı" title="The Mental Model doğrudan bağlantı" translate="no">​</a></h2>
<p>At a high level, the script wires together two related but different context
flows:</p>




















<table><thead><tr><th>Flow</th><th>AfterImage piece</th><th>Purpose</th></tr></thead><tbody><tr><td>Instruction-side context</td><td><code>QdrantDocumentProvider</code> + <code>ContextualInstructionGeneratorCallback</code></td><td>Samples documents so the simulated user asks grounded questions.</td></tr><tr><td>Response-side retrieval</td><td><code>QdrantRetriever</code> + <code>WithRAGRespondentPromptModifier</code></td><td>Retrieves relevant excerpts for the assistant before it answers.</td></tr></tbody></table>
<p>That split is the main idea. The user side gets a sampled briefing so it can ask
realistic questions. The assistant side gets retrieval results so it can answer
from the corpus instead of inventing facts.</p>
<p>The generation loop then looks like this:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">Qdrant collection</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; QdrantDocumentProvider samples context</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; ContextualInstructionGeneratorCallback creates user instructions</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; QdrantRetriever retrieves answer context for each instruction</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; WithRAGRespondentPromptModifier injects retrieved context</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; ConversationGenerator simulates user/assistant turns</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; JSONLStorage writes conversations</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="imports-the-building-blocks">Imports: The Building Blocks<a href="https://altai.dev/tr/blog/caselaw-rag-example#imports-the-building-blocks" class="hash-link" aria-label="Imports: The Building Blocks doğrudan bağlantı" title="Imports: The Building Blocks doğrudan bağlantı" translate="no">​</a></h2>
<p>The <code>afterimage</code> imports are the reusable pieces:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> afterimage </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    ConversationGenerator</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    ContextualInstructionGeneratorCallback</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    EmbeddingProviderFactory</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    GenerationMonitor</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    WithRAGRespondentPromptModifier</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> afterimage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">providers </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> QdrantDocumentProvider</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> afterimage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">retrievers </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> QdrantRetriever</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> afterimage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">storage </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> JSONLStorage</span><br></div></code></pre></div></div>
<p>You can read them as roles:</p>
<ul>
<li class=""><code>ConversationGenerator</code> is the orchestrator-facing facade. It runs the
simulated conversation and saves rows.</li>
<li class=""><code>ContextualInstructionGeneratorCallback</code> creates the initial user questions
from sampled document context.</li>
<li class=""><code>QdrantDocumentProvider</code> is a document sampler. It gives AfterImage source
text for instruction generation.</li>
<li class=""><code>QdrantRetriever</code> is a semantic retriever. It searches Qdrant for passages
relevant to the current instruction.</li>
<li class=""><code>WithRAGRespondentPromptModifier</code> adds retrieved passages to the assistant
prompt before the assistant answers.</li>
<li class=""><code>EmbeddingProviderFactory</code> creates the query embedding backend used by
<code>QdrantRetriever</code>.</li>
<li class=""><code>GenerationMonitor</code> records generation metrics and alerts.</li>
<li class=""><code>JSONLStorage</code> writes the final dataset.</li>
</ul>
<p>The Qdrant client imports are infrastructure rather than AfterImage concepts:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">from</span><span class="token plain"> qdrant_client </span><span class="token keyword" style="color:rgb(86, 156, 214)">import</span><span class="token plain"> AsyncQdrantClient</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> QdrantClient</span><br></div></code></pre></div></div>
<p>This example uses the sync client for document sampling and the async client for
retrieval.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="prompts-defining-the-two-actors">Prompts: Defining the Two Actors<a href="https://altai.dev/tr/blog/caselaw-rag-example#prompts-defining-the-two-actors" class="hash-link" aria-label="Prompts: Defining the Two Actors doğrudan bağlantı" title="Prompts: Defining the Two Actors doğrudan bağlantı" translate="no">​</a></h2>
<p>The script defines two system prompts:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">RESPONDENT_PROMPT </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token triple-quoted-string string" style="color:rgb(206, 145, 120)">"""You are a careful senior legal research assistant..."""</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">CORRESPONDENT_PROMPT </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token triple-quoted-string string" style="color:rgb(206, 145, 120)">"""You are an experienced lawyer or a legally curious client..."""</span><br></div></code></pre></div></div>
<p>In AfterImage terms:</p>
<ul>
<li class="">The <strong>respondent</strong> is the assistant being trained or simulated.</li>
<li class="">The <strong>correspondent</strong> is the user side of the conversation.</li>
</ul>
<p>The respondent prompt is strict about grounding:</p>
<ul>
<li class="">Use retrieved court opinions and excerpts.</li>
<li class="">Preserve identifiers when available.</li>
<li class="">Explain in plain English.</li>
<li class="">Say what is missing when the context is insufficient.</li>
<li class="">Do not provide real-world legal advice.</li>
</ul>
<p>The correspondent prompt shapes the simulated user:</p>
<ul>
<li class="">Ask realistic legal questions.</li>
<li class="">Stay in role.</li>
<li class="">Do not invent docket numbers or citations.</li>
<li class="">Use the provided briefing as inspiration.</li>
</ul>
<p>When adapting this example, these prompts are usually the first thing to change.
For customer support, the respondent might be a support agent and the
correspondent might be a customer. For medical education, the respondent might
be a study tutor and the correspondent might be a student. The architecture can
stay the same while the roles change.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="cli-arguments-making-the-example-reusable">CLI Arguments: Making the Example Reusable<a href="https://altai.dev/tr/blog/caselaw-rag-example#cli-arguments-making-the-example-reusable" class="hash-link" aria-label="CLI Arguments: Making the Example Reusable doğrudan bağlantı" title="CLI Arguments: Making the Example Reusable doğrudan bağlantı" translate="no">​</a></h2>
<p><code>_build_parser()</code> exposes the parts you are likely to tune:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">p</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">add_argument</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"--qdrant-url"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">p</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">add_argument</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"--collection"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">p</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">add_argument</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"--num-dialogs"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">p</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">add_argument</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"--max-turns"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">p</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">add_argument</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"--max-concurrency"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">p</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">add_argument</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"--gemini-model"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">p</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">add_argument</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"--embedding-model"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">p</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">add_argument</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string" style="color:rgb(206, 145, 120)">"--output"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>Most flags also read from environment variables. That makes the same script
usable in local demos, scheduled jobs, and documentation examples without
editing source code.</p>
<p>Important knobs:</p>













































<table><thead><tr><th>Flag</th><th>Why it matters</th></tr></thead><tbody><tr><td><code>--collection</code></td><td>Which Qdrant collection to read from.</td></tr><tr><td><code>--content-key</code></td><td>Payload field that contains document text. It must match indexing.</td></tr><tr><td><code>--max-docs</code></td><td>How many documents the document provider can sample.</td></tr><tr><td><code>--num-dialogs</code></td><td>Target number of conversations to generate.</td></tr><tr><td><code>--max-turns</code></td><td>Maximum turns per conversation. AfterImage samples uniformly from <code>1..max_turns</code>.</td></tr><tr><td><code>--max-concurrency</code></td><td>Number of concurrent conversation workers. Keep low for tight API quotas.</td></tr><tr><td><code>--gemini-max-retries</code></td><td>Retry count for transient Gemini 429/5xx responses.</td></tr><tr><td><code>--embedding-model</code></td><td>Query embedding model. It must match the indexed vector family and dimension.</td></tr><tr><td><code>--auto-improve</code></td><td>Enables evaluator-based quality retries. This costs additional LLM and embedding calls.</td></tr></tbody></table>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="qdrant-connection">Qdrant Connection<a href="https://altai.dev/tr/blog/caselaw-rag-example#qdrant-connection" class="hash-link" aria-label="Qdrant Connection doğrudan bağlantı" title="Qdrant Connection doğrudan bağlantı" translate="no">​</a></h2>
<p>The helper <code>_qdrant_kwargs()</code> builds shared client settings:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">def</span><span class="token plain"> </span><span class="token function" style="color:rgb(220, 220, 170)">_qdrant_kwargs</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">url</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"> </span><span class="token builtin" style="color:rgb(86, 156, 214)">str</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> api_key</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"> </span><span class="token builtin" style="color:rgb(86, 156, 214)">str</span><span class="token plain"> </span><span class="token operator" style="color:rgb(212, 212, 212)">|</span><span class="token plain"> </span><span class="token boolean">None</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"> </span><span class="token operator" style="color:rgb(212, 212, 212)">-</span><span class="token operator" style="color:rgb(212, 212, 212)">&gt;</span><span class="token plain"> </span><span class="token builtin" style="color:rgb(86, 156, 214)">dict</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    kwargs</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"> </span><span class="token builtin" style="color:rgb(86, 156, 214)">dict</span><span class="token plain"> </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token string" style="color:rgb(206, 145, 120)">"url"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"> url</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"> </span><span class="token string" style="color:rgb(206, 145, 120)">"timeout"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"> </span><span class="token number" style="color:rgb(181, 206, 168)">120.0</span><span class="token punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token keyword" style="color:rgb(86, 156, 214)">if</span><span class="token plain"> api_key</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">        kwargs</span><span class="token punctuation" style="color:rgb(212, 212, 212)">[</span><span class="token string" style="color:rgb(206, 145, 120)">"api_key"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">]</span><span class="token plain"> </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> api_key</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token keyword" style="color:rgb(86, 156, 214)">return</span><span class="token plain"> kwargs</span><br></div></code></pre></div></div>
<p>Both local Qdrant and Qdrant Cloud use the same shape. Cloud adds an API key;
local Docker usually does not.</p>
<p>In <code>_async_main()</code>, the script creates both clients:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">qd </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> QdrantClient</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token operator" style="color:rgb(212, 212, 212)">**</span><span class="token plain">qd_kw</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">qd_async </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> AsyncQdrantClient</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token operator" style="color:rgb(212, 212, 212)">**</span><span class="token plain">qd_kw</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>The async client is closed in <code>finally</code>:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> qd_async</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">close</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>That pattern is worth copying for any script that owns network clients.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="monitoring-and-alerts">Monitoring and Alerts<a href="https://altai.dev/tr/blog/caselaw-rag-example#monitoring-and-alerts" class="hash-link" aria-label="Monitoring and Alerts doğrudan bağlantı" title="Monitoring and Alerts doğrudan bağlantı" translate="no">​</a></h2>
<p>The monitor records metrics such as generation time, success rate, error rate,
token usage, and conversation length:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">monitor </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> GenerationMonitor</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    log_dir</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token builtin" style="color:rgb(86, 156, 214)">str</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">log_dir</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    alert_handlers</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token punctuation" style="color:rgb(212, 212, 212)">[</span><span class="token plain">on_alert</span><span class="token punctuation" style="color:rgb(212, 212, 212)">]</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    metrics_interval</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token number" style="color:rgb(181, 206, 168)">60</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>In this example, alerts are printed:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">def</span><span class="token plain"> </span><span class="token function" style="color:rgb(220, 220, 170)">on_alert</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">alert</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"> </span><span class="token operator" style="color:rgb(212, 212, 212)">-</span><span class="token operator" style="color:rgb(212, 212, 212)">&gt;</span><span class="token plain"> </span><span class="token boolean">None</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token keyword" style="color:rgb(86, 156, 214)">print</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)">f"alert - </span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token string-interpolation interpolation">alert</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token string-interpolation interpolation">name</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)"> - </span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token string-interpolation interpolation">alert</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token string-interpolation interpolation">message</span><span class="token string-interpolation interpolation punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token string-interpolation string" style="color:rgb(206, 145, 120)">"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>For production workflows, this could send alerts to a dashboard, Slack, logs, or
your own observability pipeline. For documentation and demos, printing is enough
because it shows when generation quality or provider reliability is degrading.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="instruction-side-context">Instruction-Side Context<a href="https://altai.dev/tr/blog/caselaw-rag-example#instruction-side-context" class="hash-link" aria-label="Instruction-Side Context doğrudan bağlantı" title="Instruction-Side Context doğrudan bağlantı" translate="no">​</a></h2>
<p>This block creates the document provider:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">documents </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> QdrantDocumentProvider</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    client</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">qd</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    collection_name</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">collection</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    content_key</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">content_key</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    max_docs</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">max_docs</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p><code>QdrantDocumentProvider</code> gives AfterImage documents to sample from. It is not
the same as retrieval. It is used before the conversation starts, so the
instruction generator can create grounded user questions.</p>
<p>Then the script creates the instruction callback:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">instruction_cb </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> ContextualInstructionGeneratorCallback</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    api_key</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">api_key</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    documents</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">documents</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    model_name</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">gemini_model</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    num_random_contexts</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token number" style="color:rgb(181, 206, 168)">1</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    llm_create_extras</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>This callback asks the LLM to produce user instructions from sampled context.
Its output is a <code>GeneratedInstructions</code> object containing:</p>
<ul>
<li class=""><code>instructions</code>: one or more user prompts.</li>
<li class=""><code>context</code>: the sampled source text.</li>
<li class=""><code>context_id</code> / <code>context_ids</code>: metadata for coverage tracking.</li>
</ul>
<p>In your own project, swap <code>QdrantDocumentProvider</code> for another provider if your
source material lives somewhere else:</p>
<ul>
<li class=""><code>InMemoryDocumentProvider</code> for small examples.</li>
<li class=""><code>DirectoryDocumentProvider</code> for local files.</li>
<li class=""><code>JSONLDocumentProvider</code> for prepared document rows.</li>
<li class=""><code>QdrantDocumentProvider</code> for vector database-backed corpora.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="query-embeddings">Query Embeddings<a href="https://altai.dev/tr/blog/caselaw-rag-example#query-embeddings" class="hash-link" aria-label="Query Embeddings doğrudan bağlantı" title="Query Embeddings doğrudan bağlantı" translate="no">​</a></h2>
<p>The retriever needs to embed user instructions before searching Qdrant:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">embedding_provider </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> EmbeddingProviderFactory</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">create</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">        </span><span class="token string" style="color:rgb(206, 145, 120)">"type"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"> </span><span class="token string" style="color:rgb(206, 145, 120)">"process"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">        </span><span class="token string" style="color:rgb(206, 145, 120)">"model"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"> args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">embedding_model</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">        </span><span class="token string" style="color:rgb(206, 145, 120)">"workers"</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"> args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">embedding_workers</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>This example uses a local <code>SentenceTransformer</code> process pool. That is why setup
requires:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">uv </span><span class="token function" style="color:rgb(220, 220, 170)">sync</span><span class="token plain"> </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">--extra</span><span class="token plain"> embeddings-local</span><br></div></code></pre></div></div>
<p>The important rule is that query embeddings must be compatible with indexed
vectors. Here, <code>index_corpus.py</code> stores 768-dimensional
<code>BAAI/bge-base-en-v1.5</code> vectors, so <code>generate.py</code> defaults to the same model.</p>
<p>If you index with a different model, change both indexing and querying together.
Otherwise Qdrant search may fail or return poor matches.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="response-side-retrieval">Response-Side Retrieval<a href="https://altai.dev/tr/blog/caselaw-rag-example#response-side-retrieval" class="hash-link" aria-label="Response-Side Retrieval doğrudan bağlantı" title="Response-Side Retrieval doğrudan bağlantı" translate="no">​</a></h2>
<p>The retriever is built separately from the document provider:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">retriever </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> QdrantRetriever</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    client</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">qd</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    collection_name</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">collection</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    embedding_provider</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">embedding_provider</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    async_client</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">qd_async</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    payload_key</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">content_key</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    limit</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token number" style="color:rgb(181, 206, 168)">3</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>This retrieves the top matching excerpts for a generated instruction. <code>limit=3</code>
means the assistant will see up to three retrieved passages.</p>
<p>Then the retriever is wrapped in a respondent prompt modifier:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">modifier </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> WithRAGRespondentPromptModifier</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">retriever</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">retriever</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>The modifier runs before the assistant answers. It takes the base respondent
prompt and adds retrieved context, so the assistant response is grounded in
search results.</p>
<p>This is the pattern to copy when you want RAG-style synthetic conversations:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">DocumentProvider -&gt; helps produce the user question</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">Retriever        -&gt; helps produce the assistant answer</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">PromptModifier   -&gt; injects retrieved answer context</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="storage">Storage<a href="https://altai.dev/tr/blog/caselaw-rag-example#storage" class="hash-link" aria-label="Storage doğrudan bağlantı" title="Storage doğrudan bağlantı" translate="no">​</a></h2>
<p>The generated rows are written to JSONL:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">storage </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> JSONLStorage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">conversations_path</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token builtin" style="color:rgb(86, 156, 214)">str</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">output</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>JSONL is a good default because it is easy to inspect, stream, and convert later.
Each row contains the conversation plus metadata such as sampled context,
retrieved context, and context ids.</p>
<p>If your workflow needs a database, AfterImage also supports SQL storage. The
rest of the generation composition can stay mostly the same.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-conversationgenerator">The ConversationGenerator<a href="https://altai.dev/tr/blog/caselaw-rag-example#the-conversationgenerator" class="hash-link" aria-label="The ConversationGenerator doğrudan bağlantı" title="The ConversationGenerator doğrudan bağlantı" translate="no">​</a></h2>
<p>This is where the pieces come together:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">conv_gen </span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain"> ConversationGenerator</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    respondent_prompt</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">RESPONDENT_PROMPT</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    correspondent_prompt</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">CORRESPONDENT_PROMPT</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    api_key</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">api_key</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    model_name</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">gemini_model</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    monitor</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">monitor</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    auto_improve</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">auto_improve</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    storage</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">storage</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    instruction_generator_callback</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">instruction_cb</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    respondent_prompt_modifier</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">modifier</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    embedding_provider</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">embedding_provider </span><span class="token keyword" style="color:rgb(86, 156, 214)">if</span><span class="token plain"> args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">auto_improve </span><span class="token keyword" style="color:rgb(86, 156, 214)">else</span><span class="token plain"> </span><span class="token boolean">None</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    llm_factory_kwargs</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token punctuation" style="color:rgb(212, 212, 212)">{</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token punctuation" style="color:rgb(212, 212, 212)">}</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>The most important fields are:</p>









































<table><thead><tr><th>Argument</th><th>Meaning</th></tr></thead><tbody><tr><td><code>respondent_prompt</code></td><td>System prompt for the assistant side.</td></tr><tr><td><code>correspondent_prompt</code></td><td>System prompt for the simulated user side.</td></tr><tr><td><code>instruction_generator_callback</code></td><td>Produces the first user message from sampled context.</td></tr><tr><td><code>respondent_prompt_modifier</code></td><td>Adds retrieved context before the assistant answers.</td></tr><tr><td><code>storage</code></td><td>Persists generated conversations.</td></tr><tr><td><code>monitor</code></td><td>Tracks metrics, alerts, and plots.</td></tr><tr><td><code>auto_improve</code></td><td>Runs evaluator-based retries when enabled.</td></tr><tr><td><code>llm_factory_kwargs</code></td><td>Passes provider-specific options, such as Gemini retry settings.</td></tr></tbody></table>
<p>This construction is the reusable recipe. The caselaw dataset is just one
instance of it.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="running-generation">Running Generation<a href="https://altai.dev/tr/blog/caselaw-rag-example#running-generation" class="hash-link" aria-label="Running Generation doğrudan bağlantı" title="Running Generation doğrudan bağlantı" translate="no">​</a></h2>
<p>The actual generation call is short:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> conv_gen</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">generate</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    num_dialogs</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">num_dialogs</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    max_turns</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">max_turns</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    max_concurrency</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">args</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">max_concurrency</span><span class="token punctuation" style="color:rgb(212, 212, 212)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p><code>num_dialogs</code> controls how many conversations to produce.</p>
<p><code>max_turns</code> is not "always exactly this many turns." AfterImage samples the
actual turn count uniformly from <code>1</code> through <code>max_turns</code>. With <code>max_turns=1</code>,
every conversation is single-turn.</p>
<p><code>max_concurrency</code> controls how many conversation workers run at once. Higher
values can improve throughput but also increase API pressure. For Gemini free or
tight quotas, <code>1</code> is the safest starting point.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="cleanup">Cleanup<a href="https://altai.dev/tr/blog/caselaw-rag-example#cleanup" class="hash-link" aria-label="Cleanup doğrudan bağlantı" title="Cleanup doğrudan bağlantı" translate="no">​</a></h2>
<p>The script closes resources in a <code>finally</code> block:</p>
<div class="language-python codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-python codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token keyword" style="color:rgb(86, 156, 214)">finally</span><span class="token punctuation" style="color:rgb(212, 212, 212)">:</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> embedding_provider</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">aclose</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    </span><span class="token keyword" style="color:rgb(86, 156, 214)">await</span><span class="token plain"> qd_async</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">close</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">    monitor</span><span class="token punctuation" style="color:rgb(212, 212, 212)">.</span><span class="token plain">shutdown</span><span class="token punctuation" style="color:rgb(212, 212, 212)">(</span><span class="token punctuation" style="color:rgb(212, 212, 212)">)</span><br></div></code></pre></div></div>
<p>This matters because generation scripts often run for a long time and own
process pools, async HTTP sessions, and monitoring threads.</p>
<p>If you add another resource, such as a database client or custom retriever, close
it here too.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-to-adapt-this-example">How To Adapt This Example<a href="https://altai.dev/tr/blog/caselaw-rag-example#how-to-adapt-this-example" class="hash-link" aria-label="How To Adapt This Example doğrudan bağlantı" title="How To Adapt This Example doğrudan bağlantı" translate="no">​</a></h2>
<p>Use this checklist when adapting the example to another domain:</p>
<ol>
<li class="">Replace <code>RESPONDENT_PROMPT</code> with the assistant behavior you want.</li>
<li class="">Replace <code>CORRESPONDENT_PROMPT</code> with the user role you want to simulate.</li>
<li class="">Choose a document provider for your source material.</li>
<li class="">Make sure your indexed vectors and query embedding model match.</li>
<li class="">Choose a retriever and prompt modifier if the assistant should use RAG.</li>
<li class="">Pick storage: JSONL for files, SQL for database-backed runs.</li>
<li class="">Start with low <code>num_dialogs</code> and <code>max_concurrency</code>.</li>
<li class="">Turn on <code>auto_improve</code> only when you are ready to pay for quality retries.</li>
<li class="">Inspect output JSONL before scaling up.</li>
</ol>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="common-variations">Common Variations<a href="https://altai.dev/tr/blog/caselaw-rag-example#common-variations" class="hash-link" aria-label="Common Variations doğrudan bağlantı" title="Common Variations doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="use-local-files-instead-of-qdrant-for-instruction-context">Use local files instead of Qdrant for instruction context<a href="https://altai.dev/tr/blog/caselaw-rag-example#use-local-files-instead-of-qdrant-for-instruction-context" class="hash-link" aria-label="Use local files instead of Qdrant for instruction context doğrudan bağlantı" title="Use local files instead of Qdrant for instruction context doğrudan bağlantı" translate="no">​</a></h3>
<p>If you do not need vector-backed sampling, use a file or directory provider for
the instruction side. You can still use a retriever for response-side RAG, or
skip retrieval entirely and use <code>WithContextRespondentPromptModifier</code> for simpler
context injection.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="generate-customer-support-data">Generate customer support data<a href="https://altai.dev/tr/blog/caselaw-rag-example#generate-customer-support-data" class="hash-link" aria-label="Generate customer support data doğrudan bağlantı" title="Generate customer support data doğrudan bağlantı" translate="no">​</a></h3>
<p>Keep the same shape:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">support articles -&gt; document provider -&gt; customer questions</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">support articles -&gt; retriever -&gt; agent answers</span><br></div></code></pre></div></div>
<p>Change the prompts so the respondent is a support agent and the correspondent is
a customer with realistic issues.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="generate-domain-tutoring-dialogs">Generate domain tutoring dialogs<a href="https://altai.dev/tr/blog/caselaw-rag-example#generate-domain-tutoring-dialogs" class="hash-link" aria-label="Generate domain tutoring dialogs doğrudan bağlantı" title="Generate domain tutoring dialogs doğrudan bağlantı" translate="no">​</a></h3>
<p>Use textbook sections, lecture notes, or documentation pages as source material.
The correspondent can be a beginner, advanced learner, or examiner. The
respondent can be a tutor that explains concepts and asks clarifying questions.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="skip-rag-for-simpler-datasets">Skip RAG for simpler datasets<a href="https://altai.dev/tr/blog/caselaw-rag-example#skip-rag-for-simpler-datasets" class="hash-link" aria-label="Skip RAG for simpler datasets doğrudan bağlantı" title="Skip RAG for simpler datasets doğrudan bağlantı" translate="no">​</a></h3>
<p>If every generated answer can rely on the same sampled context, you can remove
<code>QdrantRetriever</code> and <code>WithRAGRespondentPromptModifier</code>, then use a simpler
context prompt modifier. The tradeoff is less dynamic answer-time retrieval.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-this-example-teaches">What This Example Teaches<a href="https://altai.dev/tr/blog/caselaw-rag-example#what-this-example-teaches" class="hash-link" aria-label="What This Example Teaches doğrudan bağlantı" title="What This Example Teaches doğrudan bağlantı" translate="no">​</a></h2>
<p>The important lesson is not "caselaw plus Qdrant." The reusable pattern is:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">source material</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; instruction generation</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; optional retrieval</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; controlled assistant/user simulation</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">  -&gt; monitored JSONL dataset</span><br></div></code></pre></div></div>
<p>Once you understand those pieces, you can assemble your own synthetic dataset
pipeline for legal research, customer support, technical documentation, medical
education, internal knowledge bases, or any domain where conversations should be
grounded in source documents.</p>]]></content>
        <author>
            <name>Hakan Doğan</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[AfterImage is now open source for infrastructure-level dataset generation]]></title>
        <id>https://altai.dev/tr/blog/afterimage-open-source</id>
        <link href="https://altai.dev/tr/blog/afterimage-open-source"/>
        <updated>2026-04-13T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[A Python library and CLI for synthetic conversational datasets — grounded, diverse, and observable. Today we are releasing AfterImage as open source.]]></summary>
        <content type="html"><![CDATA[<blockquote>
<p>Originally published on <a href="https://medium.com/altai-dev/afterimage-is-now-open-source-for-infrastructure-level-dataset-generation-e729507c3b03" target="_blank" rel="noopener noreferrer" class="">Medium</a></p>
</blockquote>
<p>Today we are releasing AfterImage as open source software under the Apache 2.0 license, with packages on PyPI (<code>pip install afterimage</code> / <code>uv add afterimage</code>) and documentation at afterimage.altai.dev.</p>
<p>If you have ever tried to scale instruction tuning or evaluation data with an LLM, you have probably felt the same tension: raw volume is easy; useful volume is not. AfterImage exists to treat synthetic dataset generation as a systems problem — one where you can steer grounding, diversity, structure, and quality instead of hoping a clever prompt loop will "just work."</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-we-built-it">Why we built it<a href="https://altai.dev/tr/blog/afterimage-open-source#why-we-built-it" class="hash-link" aria-label="Why we built it doğrudan bağlantı" title="Why we built it doğrudan bağlantı" translate="no">​</a></h2>
<p>Modern LLM workflows lean heavily on synthetic data: bootstrapped instructions, evolved difficulty, self-curation, RAG-style corpora, and more. The research and practice around these methods make a consistent point easy to miss in day-to-day engineering: the value of synthetic data is not automatic. Without controls, you get shallow diversity, weak grounding to sources, evaluator leakage, and answers that sound right but are not anchored in anything you trust.</p>
<p>When we first designed the architecture of AfterImage internally, the core idea was simple to state and hard to operationalize: synthetic data quality has to be measured and steered, not assumed. The system was designed around explicit roles, composable stages, schema constraints where you need them, and monitoring that does not get in the way of throughput.</p>
<p>That design is still the spine of the project. What changed since that early write-up is everything around it: a YAML-driven CLI for getting JSONL in one command, export to common fine-tuning shapes, preference / DPO-style pair generation, broader provider support (including local OpenAI-compatible servers), a public docs site, and a codebase that has been exercised on real pipelines rather than described only on paper.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-afterimage-actually-does">What AfterImage actually does<a href="https://altai.dev/tr/blog/afterimage-open-source#what-afterimage-actually-does" class="hash-link" aria-label="What AfterImage actually does doğrudan bağlantı" title="What AfterImage actually does doğrudan bağlantı" translate="no">​</a></h2>
<p>AfterImage simulates conversations between two modeled roles:</p>
<ul>
<li class=""><strong>Correspondent</strong> — the side that initiates: questions, tasks, follow-ups. Behavior is driven by instruction generators and can be shaped by personas (tone, expertise, intent) so your dataset is not "one generic user, forever."</li>
<li class=""><strong>Respondent</strong> — the assistant side, defined by your system prompt and runtime prompt modifiers (for example, injecting retrieved chunks for RAG-style grounding).</li>
</ul>
<p>Generation is async-first, so you can run concurrent workers and make better use of API throughput. Persistence (JSONL or SQL) is decoupled from the core loop, so the same generation logic can land in different environments without rewrites.</p>
<p>There are two "front doors," intentionally:</p>
<ol>
<li class="">
<p><strong>CLI + config</strong> — describe a run in YAML, set keys in the environment, run <code>afterimage generate</code>. Dry-run with <code>--dry-run</code> when you want to validate the plan without spending tokens. When you are ready to train, <code>afterimage export</code> converts datasets into formats like ShareGPT-style or Hugging Face messages layouts; <code>afterimage preference</code> helps produce preference pairs for alignment workflows.</p>
</li>
<li class="">
<p><strong>Python API</strong> — compose instruction generators, respondent prompt modifiers, stopping criteria, storage, judges, and monitoring the same way the CLI does under the hood. The API is where specialized flows live: document-grounded instruction, structured extraction, tool-calling-oriented setups, custom storage, and tighter integration with your stack.</p>
</li>
</ol>
<p>If you want a quick way that your coding agents can get up and running with AfterImage, the project publishes <code>llms.txt</code> with install steps, CLI entry points, and links to the Markdown guides and examples.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="grounding-structure-and-quality-loops">Grounding, structure, and quality loops<a href="https://altai.dev/tr/blog/afterimage-open-source#grounding-structure-and-quality-loops" class="hash-link" aria-label="Grounding, structure, and quality loops doğrudan bağlantı" title="Grounding, structure, and quality loops doğrudan bağlantı" translate="no">​</a></h2>
<p>Document providers can feed local files, JSONL, in-memory lists, or Qdrant — so "what the model is allowed to know" can be a first-class input to generation, not an afterthought pasted into a prompt.</p>
<p>Structured generation uses Pydantic schemas so single-turn outputs are valid JSON-shaped objects: extraction from unstructured text, synthetic rows with typed fields, or evaluation artifacts you can consume downstream without brittle parsing.</p>
<p>Evaluation is meant to be practical at scale: embedding-based signals for fast filtering (coherence, relevance, grounding-style checks where configured) and LLM-as-judge style rubrics when you need a second pass. The framework also supports quality gates and regeneration paths (for example, <code>auto_improve</code> workflows) when you want the generator to retry or revise under explicit criteria rather than silently accepting bad rows.</p>
<p>Monitoring tracks latency, token usage, errors, and evaluation signals over time so long runs remain operable — synthetic data generation is as much an engineering workload as a modeling one.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="design-choices-that-survived-contact-with-reality">Design choices that survived contact with reality<a href="https://altai.dev/tr/blog/afterimage-open-source#design-choices-that-survived-contact-with-reality" class="hash-link" aria-label="Design choices that survived contact with reality doğrudan bağlantı" title="Design choices that survived contact with reality doğrudan bağlantı" translate="no">​</a></h2>
<p>Two ideas from the original technical framing stayed true in the shipped library:</p>
<p><strong>Composition over inheritance.</strong> You should not need to fork a monolithic "generator" class for every new corpus or policy. Behavior is injected through callbacks and strategies — instruction logic, prompt modification, storage, evaluation — so the core stays stable while your pipeline evolves.</p>
<p><strong>Provider portability.</strong> AfterImage normalizes chat sessions, structured calls, token accounting, and metadata across Gemini, OpenAI-compatible APIs (including DeepSeek and OpenRouter where applicable), and local OpenAI-compatible servers (for example vLLM, Ollama, llama.cpp). Conversation state is managed independently of any single vendor's chat API quirks.</p>
<p>Internally, recent work strengthened the parts that matter most in production: smarter context and persona sampling (including coverage-oriented behavior), clearer separation between orchestration, sampling, and quality gating, and optional capture of reasoning / thinking content from compatible providers when you want that signal preserved in your dataset.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-afterimage-is-not">What AfterImage is not<a href="https://altai.dev/tr/blog/afterimage-open-source#what-afterimage-is-not" class="hash-link" aria-label="What AfterImage is not doğrudan bağlantı" title="What AfterImage is not doğrudan bağlantı" translate="no">​</a></h2>
<p>AfterImage is not a promise that synthetic data is "as good as hand-labeled gold for all sorts of use cases" without validation, and it is not a theory of instruction learning. It is a pragmatic engine for people who need repeatable, configurable, observable generation — from a quick experiment to a large batch job with budgets and filters.</p>
<p>The same limitations that applied when we wrote the internal report still apply in public form: evaluator bias can leak into curated sets; LLM judges cost time and money; personas need care so diversity does not collapse into caricature. We ship controls and visibility so you can see those trade-offs instead of discovering them only after training.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="try-it">Try it<a href="https://altai.dev/tr/blog/afterimage-open-source#try-it" class="hash-link" aria-label="Try it doğrudan bağlantı" title="Try it doğrudan bağlantı" translate="no">​</a></h2>
<ul>
<li class=""><strong>Repository:</strong> <a href="https://github.com/altaidevorg/afterimage" target="_blank" rel="noopener noreferrer" class="">github.com/altaidevorg/afterimage</a></li>
<li class=""><strong>Documentation:</strong> afterimage.altai.dev</li>
<li class=""><strong>Install:</strong> Python 3.11+, then <code>pip install afterimage</code> or <code>uv add afterimage</code></li>
<li class=""><strong>Quick start:</strong> <code>afterimage generate -c examples/configs/basic.yaml</code> (see <code>examples/configs/</code> for RAG, local models, stopping budgets, and more)</li>
</ul>
<p>Optional extras cover local embeddings (<code>embeddings-local</code>), a small FastAPI server (<code>server</code>), and a training / demo stack (<code>training</code>) for the Gradio demo and training helpers — see <code>pyproject.toml</code> for the exact dependency sets.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="closing">Closing<a href="https://altai.dev/tr/blog/afterimage-open-source#closing" class="hash-link" aria-label="Closing doğrudan bağlantı" title="Closing doğrudan bağlantı" translate="no">​</a></h2>
<p>Open sourcing AfterImage is our invitation to treat synthetic data as infrastructure: something you configure, measure, and iterate on like any other part of the ML stack. If you build something with it — or hit a wall we should address — we would love to hear from you in issues and discussions on GitHub.</p>]]></content>
        <author>
            <name>Yusuf Sarıgöz</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Building a Small yet Highly Capable LLM-as-a-Judge: Fine-Tuning Gemma 3 4B for Evaluation Tasks]]></title>
        <id>https://altai.dev/tr/blog/llm-as-a-judge-gemma</id>
        <link href="https://altai.dev/tr/blog/llm-as-a-judge-gemma"/>
        <updated>2025-11-20T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[Open-source language models have made incredible progress in reasoning and instruction following — yet they still struggle with one crucial skill: evaluation.]]></summary>
        <content type="html"><![CDATA[<blockquote>
<p>Originally published on <a href="https://medium.com/altai-dev/building-a-small-yet-highly-capable-llm-as-a-judge-fine-tuning-gemma-3-4b-for-evaluation-tasks-4d260ef3f203" target="_blank" rel="noopener noreferrer" class="">Medium</a></p>
</blockquote>
<p>Open-source language models have made incredible progress in reasoning and instruction following — yet they still struggle with one crucial skill: evaluation. How can we train small models to judge long-form text with the nuance and consistency of giant models such as GPT-4?</p>
<p>In this project, we built a compact, open-source LLM-as-a-Judge by fine-tuning the Gemma 3 4B model on two complementary datasets from Prometheus-Eval:</p>
<ul>
<li class=""><strong>Feedback Collection</strong> → absolute scoring (1–5)</li>
<li class=""><strong>Preference Collection</strong> → pairwise ranking (A vs B)</li>
</ul>
<p>We trained two specialized models, then merged them with Nuslerp using the MergeKit library — creating a single, balanced evaluator capable of both scoring and ranking. With extensive evaluation and testing of this model, we found out that it outperforms the previously open-source SOTA model fine-tuned on this dataset despite being much smaller. We also measured the consistency among its evaluations and verified that it does a very good job without a specialized heuristic method.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-build-an-llm-as-a-judge">Why Build an LLM-as-a-Judge?<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#why-build-an-llm-as-a-judge" class="hash-link" aria-label="Why Build an LLM-as-a-Judge? doğrudan bağlantı" title="Why Build an LLM-as-a-Judge? doğrudan bağlantı" translate="no">​</a></h2>
<p>Human evaluation remains the gold standard for measuring helpfulness, factuality, and coherence in model outputs. However, it is expensive, slow, and difficult to scale. We need automated evaluations to be able to increase the number of iteration cycles and make rapid progress.</p>
<p>LLM-judges fill this gap by learning to approximate human preferences through structured feedback data. A good judge model can:</p>
<ul>
<li class="">Score model outputs on dimensions like helpfulness, factuality, etc.</li>
<li class="">Compare two responses and decide which one is better.</li>
<li class="">Provide textual rationale for its judgment.</li>
<li class="">Serve as an automated evaluation system for ongoing model training experiments.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="datasets-prometheus-eval-feedback--preference-collections">Datasets: Prometheus-Eval Feedback &amp; Preference Collections<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#datasets-prometheus-eval-feedback--preference-collections" class="hash-link" aria-label="Datasets: Prometheus-Eval Feedback &amp; Preference Collections doğrudan bağlantı" title="Datasets: Prometheus-Eval Feedback &amp; Preference Collections doğrudan bağlantı" translate="no">​</a></h2>
<p>Both datasets were released in Prometheus (2024) to teach open models fine-grained evaluation capabilities.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="feedback-collection--absolute-scoring">Feedback Collection — Absolute Scoring<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#feedback-collection--absolute-scoring" class="hash-link" aria-label="Feedback Collection — Absolute Scoring doğrudan bağlantı" title="Feedback Collection — Absolute Scoring doğrudan bağlantı" translate="no">​</a></h3>
<ul>
<li class="">1K rubrics, 20K instructions + reference answers, 100K responses + feedback</li>
<li class="">Output format — Feedback: <code>&lt;critique text&gt; [RESULT] &lt;1-5&gt;</code></li>
<li class="">Trains models to assess responses against reference answers and rubric descriptions.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="preference-collection--pairwise-judgment">Preference Collection — Pairwise Judgment<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#preference-collection--pairwise-judgment" class="hash-link" aria-label="Preference Collection — Pairwise Judgment doğrudan bağlantı" title="Preference Collection — Pairwise Judgment doğrudan bağlantı" translate="no">​</a></h3>
<ul>
<li class="">1K rubrics, 20K instructions, 200K response pairs (A/B)</li>
<li class="">Output format — Feedback: <code>&lt;comparison text&gt; [RESULT] &lt;A or B&gt;</code></li>
<li class="">Trains models to compare two responses and pick the superior one.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="dual-fine-tuning-strategy">Dual Fine-Tuning Strategy<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#dual-fine-tuning-strategy" class="hash-link" aria-label="Dual Fine-Tuning Strategy doğrudan bağlantı" title="Dual Fine-Tuning Strategy doğrudan bağlantı" translate="no">​</a></h2>
<p>We fine-tuned two separate judge models:</p>
<ul>
<li class=""><strong>Feedback model</strong> on Feedback Collection</li>
<li class=""><strong>Preference model</strong> on Preference Collection</li>
</ul>
<p>We fine-tuned attention and MLP modules of the first model using LoRA with r = 64, LoRA alpha = 32, and LoRA dropout = 0.15 on top of the feedback collection. In parallel, we fine-tuned another model on top of the preference collection with R = 128.</p>
<p>Each model captured distinct judgment behavior: one mastered calibrated scoring, the other comparative reasoning.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="merging-with-nuslerp-mergekit">Merging with Nuslerp (MergeKit)<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#merging-with-nuslerp-mergekit" class="hash-link" aria-label="Merging with Nuslerp (MergeKit) doğrudan bağlantı" title="Merging with Nuslerp (MergeKit) doğrudan bağlantı" translate="no">​</a></h3>
<p>To combine their complementary abilities, we used Nuslerp, a non-linear, spherical interpolation algorithm that merges weight spaces layer-wise for smoother semantic blending.</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">mergekit </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">--base</span><span class="token plain"> unsloth/gemma3-4b-it </span><span class="token punctuation" style="color:rgb(212, 212, 212)">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">         </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">--models</span><span class="token plain"> feedback_model preference_model </span><span class="token punctuation" style="color:rgb(212, 212, 212)">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">         </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">--weights</span><span class="token plain"> </span><span class="token number" style="color:rgb(181, 206, 168)">0.5</span><span class="token plain"> </span><span class="token number" style="color:rgb(181, 206, 168)">0.5</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(212, 212, 212)">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">         </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">--method</span><span class="token plain"> nuslerp </span><span class="token punctuation" style="color:rgb(212, 212, 212)">\</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">         </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">--output</span><span class="token plain"> gemma3-4b-judge-merged</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="evaluation-results">Evaluation Results<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#evaluation-results" class="hash-link" aria-label="Evaluation Results doğrudan bağlantı" title="Evaluation Results doğrudan bağlantı" translate="no">​</a></h2>
<p>We evaluated all models on Prometheus-Eval/Feedback-Bench and Prometheus-Eval/Preference-Bench.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="feedback-model-results">Feedback Model Results<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#feedback-model-results" class="hash-link" aria-label="Feedback Model Results doğrudan bağlantı" title="Feedback Model Results doğrudan bağlantı" translate="no">​</a></h3>
<p><img decoding="async" loading="lazy" src="https://cdn-images-1.medium.com/max/1024/1*V35oseVIO0B3N3ewbngPfg.png" alt="Feedback Model Results" class="img_ev3q"></p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="preference-model-results">Preference Model Results<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#preference-model-results" class="hash-link" aria-label="Preference Model Results doğrudan bağlantı" title="Preference Model Results doğrudan bağlantı" translate="no">​</a></h3>
<p><img decoding="async" loading="lazy" src="https://cdn-images-1.medium.com/max/1024/1*8sgNdbtHHcf3XweNXyS5Ig.png" alt="Preference Model Results" class="img_ev3q"></p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="merged-model-feedback--preference">Merged Model (Feedback + Preference)<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#merged-model-feedback--preference" class="hash-link" aria-label="Merged Model (Feedback + Preference) doğrudan bağlantı" title="Merged Model (Feedback + Preference) doğrudan bağlantı" translate="no">​</a></h3>
<p><img decoding="async" loading="lazy" src="https://cdn-images-1.medium.com/max/1024/1*nVCcSfEosmsigxfkDzhXvw.png" alt="Merged Model Results" class="img_ev3q"></p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="base-model-googlegemma3-4b-it">Base Model (google/gemma3-4b-it)<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#base-model-googlegemma3-4b-it" class="hash-link" aria-label="Base Model (google/gemma3-4b-it) doğrudan bağlantı" title="Base Model (google/gemma3-4b-it) doğrudan bağlantı" translate="no">​</a></h3>
<p><img decoding="async" loading="lazy" src="https://cdn-images-1.medium.com/max/1024/1*1Y1_IScFJ8WD5TdRPHzgGQ.png" alt="Base Model Results" class="img_ev3q"></p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="comparison-summary">Comparison Summary<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#comparison-summary" class="hash-link" aria-label="Comparison Summary doğrudan bağlantı" title="Comparison Summary doğrudan bağlantı" translate="no">​</a></h3>
<p><img decoding="async" loading="lazy" src="https://cdn-images-1.medium.com/max/1024/1*BeIcjHOT34BVVLXNTOixDw.png" alt="Comparison Summary" class="img_ev3q"></p>
<p>Our merged Gemma 3 4B-based judge matches or surpasses Prometheus 2 (8×7B) correlations while being significantly smaller.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="consistency-evaluation">Consistency Evaluation<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#consistency-evaluation" class="hash-link" aria-label="Consistency Evaluation doğrudan bağlantı" title="Consistency Evaluation doğrudan bağlantı" translate="no">​</a></h2>
<p>TrustJudge is a probabilistic evaluation framework designed to make LLM-as-a-judge systems more reliable. Traditional evaluation setups often suffer from two key issues:</p>
<ol>
<li class=""><strong>Score-Comparison Inconsistency</strong> — where lower-rated responses can outperform higher-scored ones in pairwise comparisons.</li>
<li class=""><strong>Pairwise Transitivity Inconsistency</strong> — where circular or contradictory preferences emerge (e.g., A &gt; B &gt; C &gt; A).</li>
</ol>
<p>To address these, TrustJudge introduces distribution-sensitive scoring, which converts discrete rating probabilities into continuous expectations for finer-grained judgments, and likelihood-aware aggregation, which corrects transitivity violations using bidirectional preference probabilities.</p>
<p>We experimented with the TrustJudge approach in our case. Although the method theoretically enhances evaluation consistency, in our experiments it didn't lead to any noticeable difference. This suggests that TrustJudge, despite being beneficial for zero-shot settings, may not provide additional benefits for task-specific fine-tuned models, since such models have already internalized the key evaluation signals during training.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="inconsistency-in-trustjudge">Inconsistency in TrustJudge<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#inconsistency-in-trustjudge" class="hash-link" aria-label="Inconsistency in TrustJudge doğrudan bağlantı" title="Inconsistency in TrustJudge doğrudan bağlantı" translate="no">​</a></h3>
<p>We used the prometheus-eval/Preference-Bench dataset to analyze evaluation inconsistencies. For each sample in the dataset, we evaluated two responses and selected the one with the higher score. Then, we checked whether this response had been previously preferred over the other one. If not, we marked it as an inconsistent sample.</p>
<p>After running this process across the entire dataset:</p>





















<table><thead><tr><th>Metric</th><th>Result</th></tr></thead><tbody><tr><td>Consistency</td><td>77.83%</td></tr><tr><td>Inconsistency</td><td>3.50%</td></tr><tr><td>Ties</td><td>18.67%</td></tr></tbody></table>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="limitations--next-steps">Limitations &amp; Next Steps<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#limitations--next-steps" class="hash-link" aria-label="Limitations &amp; Next Steps doğrudan bağlantı" title="Limitations &amp; Next Steps doğrudan bağlantı" translate="no">​</a></h2>
<p>Even though the model also does a decent job in non-English settings thanks to Gemma 3's multilingual capabilities, we did not run systematic benchmarks in these languages since the underlying dataset is exclusively in English. A natural next step will be to fine-tune the model explicitly for multilingual use cases.</p>
<p>The current models need reference answers to evaluate model responses. This is particularly useful to evaluate models against a validation subset where we already have a golden answer. However, another common use case is to evaluate responses supposed to be grounded on a given reference context. We will add support for such use cases as an upcoming improvement.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="links">Links<a href="https://altai.dev/tr/blog/llm-as-a-judge-gemma#links" class="hash-link" aria-label="Links doğrudan bağlantı" title="Links doğrudan bağlantı" translate="no">​</a></h2>
<ul>
<li class=""><a href="https://huggingface.co/altaidevorg" target="_blank" rel="noopener noreferrer" class="">Gemma Judge on Hugging Face</a></li>
<li class=""><a href="https://discord.gg/altai" target="_blank" rel="noopener noreferrer" class="">Altai Discord Server</a></li>
</ul>]]></content>
        <author>
            <name>Hakan Doğan</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Fine-Tuning Is Not Dead: How Synthetic QA Fine-Tuning Makes RAG Smarter]]></title>
        <id>https://altai.dev/tr/blog/synthetic-qa-fine-tuning-rag</id>
        <link href="https://altai.dev/tr/blog/synthetic-qa-fine-tuning-rag"/>
        <updated>2025-11-17T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[At Altai, we are frequently asked one question — why bother fine-tuning when you can just use RAG? Here is the answer.]]></summary>
        <content type="html"><![CDATA[<blockquote>
<p>Originally published on <a href="https://medium.com/altai-dev/fine-tuning-is-not-dead-how-synthetic-qa-fine-tuning-makes-rag-smarter-dde835a04c8c" target="_blank" rel="noopener noreferrer" class="">Medium</a></p>
</blockquote>
<p>At Altai, we're frequently asked one question:</p>
<blockquote>
<p>"Why bother fine-tuning when you can just use RAG?"</p>
</blockquote>
<p>Retrieval-Augmented Generation (RAG) — the idea of feeding external documents to a frozen large language model (LLM) — quickly became the default pattern for domain-specific enterprise AI. It's conversational, so it sounds intuitive to even non-technicals. It avoids retraining, so it's developer-friendly. It's flexible and composable, so you can build different pipelines with multiple components.</p>
<p>But something has been missing in that story: it's simply like giving a complex textbook to someone without the background and expecting expert reasoning — they can find surface-level answers to trivial questions, but they will struggle to synthesize deeper, domain-grounded insights and inevitably fall short on higher-order reasoning.</p>
<p>A recent academic paper has empirically shown this to be true. The authors of <em>The Role of Parametric Injection — A Systematic Study of Parametric Retrieval</em> systematically compared:</p>
<ul>
<li class=""><strong>Contextual RAG</strong> — standard RAG with retrieval results injected in the context</li>
<li class=""><strong>PRAG</strong> — knowledge injected in the form of LoRA adapters trained on documents</li>
<li class=""><strong>PRAG-Combine</strong> — a hybrid method combining both</li>
</ul>
<p>With this study, they concluded that RAG gets smarter, more robust, and more faithful when combined with domain-specific fine-tuning, which supports the hypothesis we built Altai around.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-core-insight-synthetic-qa-fine-tuning-improves-rag-itself">The core insight: synthetic QA fine-tuning improves RAG itself<a href="https://altai.dev/tr/blog/synthetic-qa-fine-tuning-rag#the-core-insight-synthetic-qa-fine-tuning-improves-rag-itself" class="hash-link" aria-label="The core insight: synthetic QA fine-tuning improves RAG itself doğrudan bağlantı" title="The core insight: synthetic QA fine-tuning improves RAG itself doğrudan bağlantı" translate="no">​</a></h2>
<p>The authors fine-tuned LoRA adapters (small, efficient fine-tuning modules) on document-question-answer triples — many of which were synthetically generated. These adapters captured domain-specific semantics, effectively encoding how to interpret retrieved text, not just what the text says.</p>
<p>Then they tested three setups:</p>
<ul>
<li class=""><strong>RAG</strong>: Retrieve relevant text and prompt the base model. Baseline.</li>
<li class=""><strong>PRAG</strong>: Inject fine-tuned LoRAs, no retrieval. Worse than RAG — lacks fine-grained factual details.</li>
<li class=""><strong>PRAG-Combine</strong>: Use both fine-tuned LoRAs and retrieval. Outperformed RAG on all benchmarks. Excels in the case of noisy retrieval results.</li>
</ul>
<p>The hybrid model — PRAG-Combine — consistently beat RAG across factual QA, multi-hop reasoning, and robustness tests with noisy or irrelevant passages.</p>
<p>That's the headline:</p>
<blockquote>
<p>"Fine-tuning with synthetic QA pairs doesn't compete with RAG — it amplifies and completes it exactly where it fails."</p>
</blockquote>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-this-matters-for-altai">Why this matters for Altai<a href="https://altai.dev/tr/blog/synthetic-qa-fine-tuning-rag#why-this-matters-for-altai" class="hash-link" aria-label="Why this matters for Altai doğrudan bağlantı" title="Why this matters for Altai doğrudan bağlantı" translate="no">​</a></h2>
<p>At Altai, we've built a platform where companies can train small LLMs on synthetic, domain-specific QA datasets — generated automatically from their own knowledge base.</p>
<p>The intuition has always been simple:</p>
<blockquote>
<p>"You can't expect a model that's never seen your domain to interpret your documents optimally — even if you retrieve the right ones."</p>
</blockquote>
<p>This new research provides empirical backing for that intuition. It shows that domain fine-tuning changes how the model uses retrieved context. LoRAs trained on QA pairs help the model integrate text more coherently, reduce hallucination, and stay grounded even when retrieval is messy.</p>
<p>Altai's approach — generating synthetic QAs, fine-tuning, and combining that with retrieval — aligns exactly with PRAG-Combine, the configuration that outperformed everything else.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="whats-actually-happening-under-the-hood">What's actually happening under the hood<a href="https://altai.dev/tr/blog/synthetic-qa-fine-tuning-rag#whats-actually-happening-under-the-hood" class="hash-link" aria-label="What's actually happening under the hood doğrudan bağlantı" title="What's actually happening under the hood doğrudan bağlantı" translate="no">​</a></h2>
<p>The paper's layer-wise analysis revealed the mechanism that explains why PRAG-Combine works better: Fine-tuned adapters increase parametric knowledge scores in the later layers of the transformer — the parts responsible for semantic reasoning.</p>
<p>In plain terms, fine-tuning on synthetic QA pairs teaches the model "how to think" about domain-specific information.</p>
<p>That means when the retriever feeds a new document, the model doesn't just read it literally — it interprets it through the lens of its domain knowledge. It knows what really matters, resolves conflicts, and synthesizes a thoughtful response instead of copy-pasting from the context.</p>
<p>In the Altai pipeline, this is exactly what happens:</p>
<ol>
<li class=""><strong>Synthetic QA generation</strong>: Altai's Afterimage framework leveraging our custom-made models generates diverse, domain-relevant questions and answers from a company's documents.</li>
<li class=""><strong>Fine-tuning adapters</strong>: These QA pairs train a smaller model or LoRA module to internalize the domain's structure — concepts, relationships, typical instruction forms.</li>
<li class=""><strong>Hybrid inference (letsearch)</strong>: During serving, we still use retrieval — but now the model interprets the retrieved chunks with domain intuition.</li>
</ol>
<p>Result: faster convergence, better grounding, and much more reliable outputs in specialized domains (law, healthcare, finance, etc.).</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-synthetic-qa-data-works-so-well">Why synthetic QA data works so well<a href="https://altai.dev/tr/blog/synthetic-qa-fine-tuning-rag#why-synthetic-qa-data-works-so-well" class="hash-link" aria-label="Why synthetic QA data works so well doğrudan bağlantı" title="Why synthetic QA data works so well doğrudan bağlantı" translate="no">​</a></h2>
<p>Synthetic QA pairs are not just convenient — they're structurally ideal training data:</p>
<ul>
<li class=""><strong>Question form</strong> teaches information need recognition — what matters in the text.</li>
<li class=""><strong>Answer form</strong> teaches fact grounding — where and how to extract it.</li>
<li class="">Together, they encode "retrieval intent" directly into the model's parameters.</li>
</ul>
<p>The paper shows that even when generated automatically, such QA data improves interpretive alignment between the model and retrieved context.</p>
<p>In other words, the model learns not just facts, but also how to think in the light of facts.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-big-picture-fine-tuning-and-retrieval-arent-rivals">The big picture: fine-tuning and retrieval aren't rivals<a href="https://altai.dev/tr/blog/synthetic-qa-fine-tuning-rag#the-big-picture-fine-tuning-and-retrieval-arent-rivals" class="hash-link" aria-label="The big picture: fine-tuning and retrieval aren't rivals doğrudan bağlantı" title="The big picture: fine-tuning and retrieval aren't rivals doğrudan bağlantı" translate="no">​</a></h2>
<p>The industry treated RAG as the "fine-tuning killer." However, this paper demonstrates that RAG does not suffice in settings involving domain knowledge. The solution isn't RAG or fine-tuning — it's RAG plus domain fine-tuning via synthetic data.</p>
<p>Altai operationalizes that theoretical finding. Our platform:</p>
<ul>
<li class=""><strong>Generates synthetic QA datasets</strong> tailored to your documents thanks to our enterprise-grade synthetic dataset engine (Afterimage).</li>
<li class=""><strong>Fine-tunes or LoRA-adapts models</strong> to embed domain semantics on the Altai platform.</li>
<li class=""><strong>Combines that with retrieval</strong> through our RAG layer (letsearch) for deployment.</li>
</ul>
<p>The result: leaner models that know your domain and interpret your knowledge base intelligently. Retrieval gets you access to knowledge, but fine-tuning teaches you how to use it.</p>
<p>Altai's synthetic QA fine-tuning turns your documents into domain intuition — the missing half of the retrieval story.</p>]]></content>
        <author>
            <name>Yusuf Sarıgöz</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Small Models, Big Impact: Why Altai's SLMs Outperform LLMs for Business Needs]]></title>
        <id>https://altai.dev/tr/blog/small-models-big-impact</id>
        <link href="https://altai.dev/tr/blog/small-models-big-impact"/>
        <updated>2025-07-03T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[Businesses are eager to harness AI — but the biggest, most expensive models aren't always the right answer. Here's why smaller, domain-specific models win.]]></summary>
        <content type="html"><![CDATA[<blockquote>
<p>Originally published on <a href="https://medium.com/altai-dev/small-models-big-impact-why-altais-slms-outperform-llms-for-business-needs-ae4de5ed6b42" target="_blank" rel="noopener noreferrer" class="">Medium</a></p>
</blockquote>
<p>In today's rapidly evolving digital landscape, businesses of all sizes are eager to harness the transformative power of Artificial Intelligence (AI). Large Language Models (LLMs) like GPT-4 or Gemini have captured headlines with their broad capabilities, and for good reason. However, the initial fascination with these massive, general-purpose models is giving way to a more nuanced understanding of their practical limitations for many enterprise applications. This is where Altai's Small Language Models (SLMs) emerge as a game-changer, offering a pragmatic and powerful alternative designed specifically for business needs.</p>
<p>Even tech giants like NVIDIA acknowledge this shift. As a recent research paper from NVIDIA emphasizes, "SLMs are sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems," suggesting that small models — not large — will be the backbone of scalable and sustainable AI systems in the future (Belcak et al., 2025).</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-problem-with-bigger-is-better-why-llms-fall-short-for-businesses">The Problem with "Bigger is Better": Why LLMs Fall Short for Businesses<a href="https://altai.dev/tr/blog/small-models-big-impact#the-problem-with-bigger-is-better-why-llms-fall-short-for-businesses" class="hash-link" aria-label="The Problem with &quot;Bigger is Better&quot;: Why LLMs Fall Short for Businesses doğrudan bağlantı" title="The Problem with &quot;Bigger is Better&quot;: Why LLMs Fall Short for Businesses doğrudan bağlantı" translate="no">​</a></h2>
<p>While LLMs are remarkable in their versatility and extensive linguistic knowledge, their inherent design often presents significant drawbacks for corporate adoption:</p>
<ul>
<li class="">
<p><strong>Prohibitive Costs:</strong> Training and operating LLMs are incredibly expensive due to their immense scale and computational demands. They require high-end GPUs, large infrastructure, and consume substantial energy, leading to high cloud computing bills and significant ongoing operational expenses. For instance, training a model like GPT-3 could cost approximately $1.4 million per session, and per-query costs can be as high as $0.09 compared to less than $0.0004 for SLMs.</p>
</li>
<li class="">
<p><strong>Inefficiency for Specific Tasks:</strong> LLMs are generalists, designed to emulate broad human intelligence across diverse domains. While versatile, this broadness can paradoxically make them less precise or even lead to "hallucinations" — generating factually incorrect or irrelevant information — when applied to highly specialized fields without specific adaptation. This makes them less reliable for critical business functions like legal analysis or medical diagnostics where precision is paramount.</p>
</li>
<li class="">
<p><strong>Significant Data Security Concerns:</strong> Most leading LLMs (like GPT, Claude, Gemini) are cloud-based and process user data through third-party providers. This poses substantial data security and privacy risks for businesses handling sensitive or regulated information, as data leaves the company's control. Compliance with regulations like GDPR or HIPAA becomes a complex challenge.</p>
</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="altais-right-sized-ai-solution-the-power-of-domain-specific-slms">Altai's "Right-Sized AI" Solution: The Power of Domain-Specific SLMs<a href="https://altai.dev/tr/blog/small-models-big-impact#altais-right-sized-ai-solution-the-power-of-domain-specific-slms" class="hash-link" aria-label="Altai's &quot;Right-Sized AI&quot; Solution: The Power of Domain-Specific SLMs doğrudan bağlantı" title="Altai's &quot;Right-Sized AI&quot; Solution: The Power of Domain-Specific SLMs doğrudan bağlantı" translate="no">​</a></h2>
<p>Altai directly addresses these challenges by focusing on Small Language Models (SLMs), offering a unique and timely solution for businesses. Altai's platform is designed to democratize AI access by empowering businesses to create and train their own custom SLMs through a revolutionary approach.</p>
<p>Here's how Altai's SLMs deliver a big impact:</p>
<ul>
<li class="">
<p><strong>On-Premise (Company-Internal) and Secure Deployment:</strong> Altai understands the critical need for data sovereignty and privacy. Unlike many cloud-based LLMs, Altai allows for on-premise deployment, ensuring sensitive corporate data remains securely within the company's own infrastructure. This mitigates the risks of data leakage and simplifies compliance with strict regulations like GDPR, KVKK, and HIPAA.</p>
</li>
<li class="">
<p><strong>Superior Cost-Efficiency:</strong> Altai's focus on SLMs translates into significantly lower costs across the entire AI lifecycle.</p>
<ul>
<li class=""><strong>Reduced Infrastructure &amp; Energy:</strong> SLMs require substantially less computational power and memory for training and inference. This means lower hardware investments, reduced cloud bills (or no cloud bills with on-premise deployment), and a smaller carbon footprint, consuming up to 60% less energy than LLMs.</li>
<li class=""><strong>Faster &amp; Cheaper Training/Fine-tuning:</strong> Altai's models can be trained and fine-tuned rapidly, in weeks, days, or even hours — a stark contrast to the months often required for LLMs. This dramatically reduces development costs.</li>
<li class=""><strong>Lower Operational Costs per Query:</strong> For high-volume tasks, the operational cost per query for an SLM can be over 100 times lower than for an LLM. This makes Altai's solutions highly budget-friendly for businesses with many AI interactions.</li>
</ul>
</li>
<li class="">
<p><strong>Higher Accuracy in Targeted Tasks &amp; Reduced Hallucinations:</strong> Altai develops domain-specific SLMs that are trained primarily on enterprise-specific documents and synthetically generated data derived from these documents.</p>
<ul>
<li class="">This focused training on curated datasets allows SLMs to deeply understand the nuances, terminology, and factual landscape of a specific field. This leads to superior accuracy and reliability in targeted tasks compared to general-purpose LLMs.</li>
<li class="">This approach inherently mitigates hallucinations, as the model's knowledge scope is narrowed and grounded in verified company data, reducing the likelihood of generating false or irrelevant information.</li>
<li class="">Altai also leverages techniques like <strong>Retrieval-Augmented Generation (RAG)</strong> to provide SLMs with real-time access to up-to-date, transparent, and verifiable external or internal knowledge bases, further enhancing factual consistency and reducing hallucinations.</li>
</ul>
</li>
<li class="">
<p><strong>Faster Inference Times:</strong> The smaller size and streamlined architecture of SLMs result in significantly faster response times. This low latency is crucial for real-time applications like customer service chatbots or interactive systems, enhancing user experience and operational fluidity.</p>
</li>
<li class="">
<p><strong>User-Friendly No-Code Platform:</strong> Altai offers an intuitive no-code interface. This democratizes AI creation, allowing businesses, including SMEs and non-technical users, to easily create and train custom AI-powered language models by simply uploading their documents, eliminating the need for specialized coding knowledge.</p>
</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="conclusion">Conclusion<a href="https://altai.dev/tr/blog/small-models-big-impact#conclusion" class="hash-link" aria-label="Conclusion doğrudan bağlantı" title="Conclusion doğrudan bağlantı" translate="no">​</a></h2>
<p>The shift towards "right-sized AI" solutions is gaining significant traction in the enterprise world. Altai's focus on domain-specific SLMs directly addresses the critical pain points of high costs, inefficiency, and data security that plague traditional LLM deployments. By offering cost-efficient, high-accuracy, fast-inference, and secure on-premise solutions through a user-friendly no-code platform, Altai empowers businesses of all sizes to unlock the full, transformative potential of AI. This strategic approach not only enhances operational efficiency and reduces errors but also secures a tangible competitive advantage in today's dynamic market.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="citation">Citation<a href="https://altai.dev/tr/blog/small-models-big-impact#citation" class="hash-link" aria-label="Citation doğrudan bağlantı" title="Citation doğrudan bağlantı" translate="no">​</a></h2>
<p>Belcak, P., Heinrich, G., Diao, S., Fu, Y., Dong, X., Muralidharan, S., Lin, Y. C., &amp; Molchanov, P. (2025).
Small Language Models are the Future of Agentic AI. arXiv preprint.
<a href="https://arxiv.org/abs/2506.02153" target="_blank" rel="noopener noreferrer" class="">https://arxiv.org/abs/2506.02153</a></p>]]></content>
        <author>
            <name>Efe Can Çeliksoy</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Launching Altai: Unlocking Enterprise AI with Customization & Open Source]]></title>
        <id>https://altai.dev/tr/blog/launching-altai</id>
        <link href="https://altai.dev/tr/blog/launching-altai"/>
        <updated>2025-05-29T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[Enterprises are eager to leverage LLMs but adoption is hard. Here's why we built Altai — and what we're doing differently.]]></summary>
        <content type="html"><![CDATA[<blockquote>
<p>Originally published on <a href="https://medium.com/altai-dev/launching-altaiunlocking-enterprise-ai-with-customization-open-source-0238d43f27c5" target="_blank" rel="noopener noreferrer" class="">Medium</a></p>
</blockquote>
<p>Enterprises are eager to leverage large language models (LLMs) to drive automation, enhance decision-making, and streamline operations. However, adopting AI at scale comes with significant hurdles:</p>
<ul>
<li class=""><strong>Customization</strong>: Off-the-shelf models don't fit specific enterprise needs.</li>
<li class=""><strong>Data Privacy</strong>: Cloud-hosted solutions pose compliance risks.</li>
<li class=""><strong>Talent Gap</strong>: Hiring AI experts is costly and competitive.</li>
</ul>
<p>At Altai, we're solving these challenges by providing enterprises with a no-code AI platform for effortless LLM customization and deployment, whether in the cloud or on-premises. Our vision is to democratize AI by combining the power of open-source models with high-quality synthetic dataset generation and enterprise-grade adaptation.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-open-source-ai-opportunity">The Open-Source AI Opportunity<a href="https://altai.dev/tr/blog/launching-altai#the-open-source-ai-opportunity" class="hash-link" aria-label="The Open-Source AI Opportunity doğrudan bağlantı" title="The Open-Source AI Opportunity doğrudan bağlantı" translate="no">​</a></h2>
<p>Open-source LLMs like Llama and DeepSeek have made cutting-edge AI more accessible. However, enterprises require tailored solutions — domain-specific fine-tuning, secure on-prem deployment, and reliable data pipelines. As Y Combinator recently stated:</p>
<blockquote>
<p>"Just as RedHat and YC W15 company GitLab built public companies around open-source projects, we are looking for companies built around open source AI models like DeepSeek and Llama. There is a huge opportunity to build startups that offer support and services to help people use open-source AI."</p>
</blockquote>
<p>Altai seizes this opportunity by bridging the gap between open-source AI and real-world enterprise needs.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-altai-works">How Altai Works<a href="https://altai.dev/tr/blog/launching-altai#how-altai-works" class="hash-link" aria-label="How Altai Works doğrudan bağlantı" title="How Altai Works doğrudan bağlantı" translate="no">​</a></h2>
<p>Our end-to-end workflow enables organizations to customize and deploy AI models without writing code:</p>
<ol>
<li class=""><strong>Upload Documents</strong> → Process business-specific data.</li>
<li class=""><strong>Fine-Tune or Continual Pretrain</strong> → Adapt smaller LLMs to fit unique requirements.</li>
<li class=""><strong>Generate Synthetic Datasets</strong> → Use purpose-built LLMs and sophisticated algorithms to create and validate high-quality training data.</li>
<li class=""><strong>Instruction Fine-Tuning</strong> → Enhance models for enterprise-specific use cases by fine-tuning them on the validated synthetic dataset.</li>
<li class=""><strong>Deploy Securely</strong> → Deploy models on-prem or in the cloud.</li>
</ol>
<p>By integrating automation at every step, Altai empowers organizations to build smarter, more relevant AI solutions in days rather than months.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-now">Why Now?<a href="https://altai.dev/tr/blog/launching-altai#why-now" class="hash-link" aria-label="Why Now? doğrudan bağlantı" title="Why Now? doğrudan bağlantı" translate="no">​</a></h2>
<p>The demand for enterprise-ready AI is exploding. Companies struggle with:</p>
<ul>
<li class=""><strong>The need for domain-specific customization</strong> to get real business value from LLMs.</li>
<li class=""><strong>Vendor lock-in</strong> from closed-source AI providers.</li>
<li class=""><strong>Skyrocketing compute costs</strong> for training and inference.</li>
</ul>
<p>Altai eliminates these barriers, making AI useful, adaptable, enterprise-ready, and cost-effective.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-future-powered-by-open-customizable-ai">A Future Powered by Open, Customizable AI<a href="https://altai.dev/tr/blog/launching-altai#a-future-powered-by-open-customizable-ai" class="hash-link" aria-label="A Future Powered by Open, Customizable AI doğrudan bağlantı" title="A Future Powered by Open, Customizable AI doğrudan bağlantı" translate="no">​</a></h2>
<p>The AI landscape is shifting. Enterprises no longer need to rely on closed, black-box models from a handful of providers. With Altai, we're giving companies the tools to take control of their AI strategy — leveraging open-source innovation while ensuring customization, security, and scalability.</p>
<p><strong>The future of AI is open</strong>, and Altai is here to make it work for you.</p>
<p>We're just getting started. If you're building with open-source AI and looking for enterprise-ready solutions, let's talk.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-road-ahead">The Road Ahead<a href="https://altai.dev/tr/blog/launching-altai#the-road-ahead" class="hash-link" aria-label="The Road Ahead doğrudan bağlantı" title="The Road Ahead doğrudan bağlantı" translate="no">​</a></h2>
<p>We're building the foundation for scalable enterprise AI — a no-code, high-performance AI platform that empowers companies to own and control their models. If you're an enterprise looking to deploy AI without hiring a massive ML team, Altai is your answer.</p>
<p><a href="https://altai.dev/" target="_blank" rel="noopener noreferrer" class="">Sign up for early access</a></p>]]></content>
        <author>
            <name>Yusuf Sarıgöz</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Distilling Efficiency: Experiments in Compressing BAAI/bge-m3 using a Synthetic Dataset]]></title>
        <id>https://altai.dev/tr/blog/distilling-bge-m3</id>
        <link href="https://altai.dev/tr/blog/distilling-bge-m3"/>
        <updated>2025-05-16T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[The rapid progress in NLP is mainly due to large neural network models — but size comes at a cost. Here's our approach to compressing bge-m3 with synthetic data.]]></summary>
        <content type="html"><![CDATA[<blockquote>
<p>Originally published on <a href="https://medium.com/altai-dev/distilling-efficiency-experiments-in-compressing-baai-bge-m3-using-a-synthetic-dataset-9430e21c6b8f" target="_blank" rel="noopener noreferrer" class="">Medium</a></p>
</blockquote>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="introduction">Introduction<a href="https://altai.dev/tr/blog/distilling-bge-m3#introduction" class="hash-link" aria-label="Introduction doğrudan bağlantı" title="Introduction doğrudan bağlantı" translate="no">​</a></h2>
<p>The rapid progress in Natural Language Processing (NLP) is mainly due to the development of large and complex neural network models. While these models perform very well, their size and high computational needs make them hard to use in real-world applications, especially for specific types of text. Many use cases — like semantic search and analyzing synthetic questions — could benefit from AI tools, but large models are often too costly and slow.</p>
<p>This study explores distilling the powerful multilingual BAAI/bge-m3 text embedding model into smaller versions with 8, 6, and 4 layers. The goal is to find the right balance between model size, speed, and retrieval performance using a synthetic dataset of contextual texts and generated questions. As NLP models become more complex, distillation helps make advanced AI more accessible and practical, especially for resource-limited needs.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-background">1. Background<a href="https://altai.dev/tr/blog/distilling-bge-m3#1-background" class="hash-link" aria-label="1. Background doğrudan bağlantı" title="1. Background doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="11-what-is-model-distillation">1.1 What is Model Distillation?<a href="https://altai.dev/tr/blog/distilling-bge-m3#11-what-is-model-distillation" class="hash-link" aria-label="1.1 What is Model Distillation? doğrudan bağlantı" title="1.1 What is Model Distillation? doğrudan bağlantı" translate="no">​</a></h3>
<p>Model distillation — also known as knowledge distillation — is a technique used to compress large, powerful models by training smaller models to replicate their behavior. First introduced by Hinton et al. (2015), this approach allows a smaller "student" model to approximate the performance of a larger "teacher" model while using significantly fewer parameters. The result is faster inference, lower memory requirements, and reduced deployment costs — ideal for real-world applications that require scalability and efficiency.</p>
<p>At the heart of this process is the idea that the teacher model has already learned rich patterns and decision boundaries from extensive training data. Instead of training the student model solely on standard "hard" labels (like class IDs), distillation trains it on the "soft targets" produced by the teacher model. These soft targets — such as the probability distributions generated at the output layer — convey more nuanced information about how the teacher interprets the data. This helps the student not just learn the correct answers, but also gain insight into the reasoning behind the teacher's decisions.</p>
<p>In this work, the BAAI/bge-m3 model serves as the teacher. It is a state-of-the-art text embedding model, known for its effectiveness and versatility in a wide range of retrieval and representation tasks. It stands out in three important dimensions:</p>
<ul>
<li class=""><strong>Multi-Functionality</strong>: BAAI/bge-m3 supports dense retrieval, sparse retrieval (based on lexical matching), and multi-vector retrieval — all within a single model architecture.</li>
<li class=""><strong>Multi-Linguality</strong>: It works across more than 100 languages, mapping them into a shared semantic space to enable both monolingual and cross-lingual search.</li>
<li class=""><strong>Multi-Granularity</strong>: The model handles input texts of varying lengths, from short phrases to full-length documents — up to 8192 tokens — making it suitable for diverse use cases.</li>
</ul>
<p>By distilling this teacher into smaller variants, the goal is to retain as much of its semantic power as possible, while greatly improving efficiency for deployment in real-world systems.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-distilled-variants">2. Distilled Variants<a href="https://altai.dev/tr/blog/distilling-bge-m3#2-distilled-variants" class="hash-link" aria-label="2. Distilled Variants doğrudan bağlantı" title="2. Distilled Variants doğrudan bağlantı" translate="no">​</a></h2>
<p>The primary goal of this work was to create smaller, faster versions of the BAAI/bge-m3 model while retaining as much of its semantic understanding capabilities as possible, particularly for the contextual texts within the synthetic dataset. The motivation stems from the observation that while powerful models are highly effective, their large size can make them prohibitively expensive and slow for large-scale, low-latency applications.</p>
<p>The student models are:</p>
<ul>
<li class=""><code>altaidevorg/bge-m3-distill-8l</code> (8-layer distilled model)</li>
<li class=""><code>altaidevorg/bge-m3-distill-6l</code> (6-layer distilled model)</li>
<li class=""><code>altaidevorg/bge-m3-distill-4l</code> (4-layer distilled model)</li>
</ul>
<p>The student models were derived from the BAAI/bge-m3 teacher model by reducing the number of model layers. This is a common and effective approach to creating smaller versions of transformer-based models.</p>
<p><strong>Architecture:</strong> Transformer encoder with 4, 6, and 8 layers respectively for the distill-4l, distill-6l, and distill-8l models. The 8-layer model has a 366M parameter size, a significant reduction from the original 24-layer teacher. Other architectural parameters (hidden size, number of attention heads) are kept consistent with the layer configuration of the teacher model.</p>
<p><strong>Training Data:</strong> Distillation used a mix of general text corpora and domain-specific contextual texts from the synthetic dataset (including question-answer pairs). This combination aimed to ensure both general understanding and adaptation to the target domain.</p>
<p><strong>Use Cases:</strong> These models are intended for:</p>
<ul>
<li class="">Semantic search</li>
<li class="">Information retrieval</li>
<li class="">Document clustering</li>
<li class="">Embedding components in RAG systems, especially for contextual texts in the synthetic dataset</li>
</ul>
<p><strong>Performance:</strong> The 8-layer model processed about 454 texts/sec on a T4 GPU — 2.5x faster than the teacher model (175 texts/sec). It also showed strong alignment with the teacher's output, scoring a Spearman Cosine of 0.965 and MSE of 0.006 on the test set.</p>
<p><strong>Limitations:</strong></p>
<ul>
<li class="">May lose some depth on complex tasks compared to the full teacher model.</li>
<li class="">Might generalize less effectively to unrelated domains or unsupported languages.</li>
<li class="">Could inherit biases from the teacher model or training data.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-experimental-setup">3. Experimental Setup<a href="https://altai.dev/tr/blog/distilling-bge-m3#3-experimental-setup" class="hash-link" aria-label="3. Experimental Setup doğrudan bağlantı" title="3. Experimental Setup doğrudan bağlantı" translate="no">​</a></h2>
<p>To evaluate the performance of the distilled models, a robust experimental setup was designed, focusing on a semantic search task using a synthetic dataset composed of contextual texts and synthetically generated questions.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="31-dataset-and-vector-search-process">3.1 Dataset and Vector Search Process<a href="https://altai.dev/tr/blog/distilling-bge-m3#31-dataset-and-vector-search-process" class="hash-link" aria-label="3.1 Dataset and Vector Search Process doğrudan bağlantı" title="3.1 Dataset and Vector Search Process doğrudan bağlantı" translate="no">​</a></h3>
<p>The dataset used in the evaluation process is a set of synthetically generated questions derived from contextual texts within the synthetic dataset. Each query is converted into a high-dimensional vector representation that captures its semantic meaning. This embedding is then used to search against the pre-indexed document embeddings in the Qdrant vector database. The system retrieves the top-k most semantically similar document IDs, where k is varied across multiple evaluation rounds (10, 20, 30, 40, 50).</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="32-evaluation-metric-precisionk">3.2 Evaluation Metric: Precision@k<a href="https://altai.dev/tr/blog/distilling-bge-m3#32-evaluation-metric-precisionk" class="hash-link" aria-label="3.2 Evaluation Metric: Precision@k doğrudan bağlantı" title="3.2 Evaluation Metric: Precision@k doğrudan bağlantı" translate="no">​</a></h3>
<p>Precision@k measures the proportion of overlapping items found within the top k results returned by two different models for the same query. When comparing Model A and Model B:</p>
<p>$$P@k_{A vs. B} = \frac{|{A\text{'s top-k docs}} \cap {B\text{'s top-k docs}}|}{k}$$</p>
<p>In this experimental context, Precision@k is used to compare the student models against each other (4L vs. 6L, 4L vs. 8L, 6L vs. 8L) for k values of 10, 20, 30, 40, and 50.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="4-results">4. Results<a href="https://altai.dev/tr/blog/distilling-bge-m3#4-results" class="hash-link" aria-label="4. Results doğrudan bağlantı" title="4. Results doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="41-performance-precisionk-scores">4.1 Performance (Precision@k Scores)<a href="https://altai.dev/tr/blog/distilling-bge-m3#41-performance-precisionk-scores" class="hash-link" aria-label="4.1 Performance (Precision@k Scores) doğrudan bağlantı" title="4.1 Performance (Precision@k Scores) doğrudan bağlantı" translate="no">​</a></h3>
<p><strong>Observations:</strong></p>
<ul>
<li class="">
<p>The <strong>bge-m3-distill-6l</strong> and <strong>bge-m3-distill-8l</strong> models exhibit the highest similarity in their retrieval behavior across all tested values of k. P@k scores range from 0.700 at k=10 to 0.792 at k=50, indicating that their top retrieved documents are largely overlapping. This suggests that, despite having fewer layers, the 6-layer model closely mirrors the retrieval patterns of the more expressive 8-layer variant.</p>
</li>
<li class="">
<p>In contrast, the <strong>bge-m3-distill-4l</strong> model shows noticeably lower alignment with both the 6L and 8L models. The P@10 score between the 4L and 8L models is 0.580, which gradually increases to 0.665 at P@50. This indicates that the 4-layer model diverges more in selecting the most relevant documents, particularly in the top-ranked results.</p>
</li>
<li class="">
<p>A broader trend observed across all inter-model comparisons is that P@k values generally increase or stabilize as k increases. This means that while the most highly ranked results may vary more between models — especially between those with greater architectural differences like 4L vs. 8L — the overlap in retrieved documents grows when a larger set of top results is considered.</p>
</li>
</ul>
<blockquote>
<p><strong>Efficiency Gains:</strong> The bge-m3-distill-8l model provides an empirically measured inference speedup of approximately <strong>2.5x</strong> compared to the original 24-layer BAAI/bge-m3 teacher model, alongside substantial reductions in model size (366M parameters). The 6L and 4L variants are projected to offer even greater efficiency, with estimated speedups of approximately <strong>3.1x</strong> and <strong>4.0x</strong> respectively.</p>
</blockquote>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="conclusion">Conclusion<a href="https://altai.dev/tr/blog/distilling-bge-m3#conclusion" class="hash-link" aria-label="Conclusion doğrudan bağlantı" title="Conclusion doğrudan bağlantı" translate="no">​</a></h2>
<p>This study on distilling the BAAI/bge-m3 model has resulted in the development of a series of more lightweight and computationally efficient embedding models. These distilled variants were designed to strike a balance between semantic performance and practical deployment needs such as speed and resource efficiency.</p>
<p>The 8-layer model emerges as a robust baseline, offering a 2.5x speedup over the original 24-layer teacher model with minimal loss in retrieval quality. It serves as a strong default option for most retrieval-based applications.</p>
<p>The 6-layer model demonstrates impressive alignment with the 8-layer variant, achieving a P@50 of 0.792, indicating that it preserves much of the retrieval behavior while being even more resource-efficient. This makes it an attractive choice for environments where performance needs to be balanced with greater throughput or tighter latency constraints.</p>
<p>The 4-layer model, while showing the largest deviation in top-k retrieval overlap (P@10 of 0.580 compared to the 8-layer model), offers the highest projected speedup of approximately 4.0x. Its lightweight architecture makes it especially well-suited for use cases where maximizing inference speed and minimizing computational footprint are paramount, even at the cost of some retrieval precision.</p>
<p>Together, these models provide a flexible set of options for embedding-based applications, enabling practitioners to make informed decisions based on their specific trade-off requirements between speed, scale, and retrieval fidelity.</p>]]></content>
        <author>
            <name>Aleynaahukmet</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Are you tired of getting only garbage instead of Markdown from PDFs? Yeah, same.]]></title>
        <id>https://altai.dev/tr/blog/llm-food-pdf-to-markdown</id>
        <link href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown"/>
        <updated>2025-05-16T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[That's why we built llm-food — a FastAPI-based service that converts documents and URLs into clean, LLM-friendly Markdown with batch processing support.]]></summary>
        <content type="html"><![CDATA[<blockquote>
<p>Originally published on <a href="https://medium.com/altai-dev/are-you-tired-of-getting-only-garbage-instead-of-markdown-from-pdfs-yeah-same-f7816fe2f87b" target="_blank" rel="noopener noreferrer" class="">Medium</a></p>
</blockquote>
<p>That's why we built <strong>llm-food</strong> — a FastAPI-based service that converts documents and URLs into clean, LLM-friendly Markdown. It supports batch processing, integrates with Google's Gemini for blazing-fast PDF OCR, and wraps everything in a CLI and Python client that doesn't make you cry.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="tldr">tl;dr<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#tldr" class="hash-link" aria-label="tl;dr doğrudan bağlantı" title="tl;dr doğrudan bağlantı" translate="no">​</a></h2>
<ul>
<li class="">Converts PDFs, DOCX, RTF, PPTX, and HTML/web pages into Markdown.</li>
<li class="">Synchronous and async modes via HTTP API, CLI, and Python client.</li>
<li class="">Batch PDFs go through <strong>Gemini Batch Prediction API</strong> for high throughput ($1 for ~6,000 pages).</li>
<li class="">Tracks batch jobs in DuckDB.</li>
<li class="">Fully dockerized. Optional auth. One-liner CLI.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why">Why?<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#why" class="hash-link" aria-label="Why? doğrudan bağlantı" title="Why? doğrudan bağlantı" translate="no">​</a></h2>
<p>Extracting clean text from PDFs is still a mess.</p>
<p>You've probably seen dockling, marker, or pymupdf4llm. They're okay — but they're either slow, resource-hungry, or AGPL-licensed (ouch!). If you just want clean Markdown to fine-tune or RAG with your LLM, it's way more effort than it should be.</p>
<p>Enter <strong>Gemini Batch Prediction</strong>. It's fast and insanely cheap — $1 per 6,000 pages. But it's not dev-friendly.</p>
<p>So we wrapped it all up in a neat little microservice: llm-food.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-you-get">What You Get<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#what-you-get" class="hash-link" aria-label="What You Get doğrudan bağlantı" title="What You Get doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="features">Features<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#features" class="hash-link" aria-label="Features doğrudan bağlantı" title="Features doğrudan bağlantı" translate="no">​</a></h3>
<ul>
<li class="">Convert:
<ul>
<li class="">PDFs (Gemini, pymupdf4llm, or pypdf)</li>
<li class="">DOC/DOCX (via mammoth)</li>
<li class="">RTF (via striprtf)</li>
<li class="">PPTX (python-pptx)</li>
<li class="">HTML or URLs (trafilatura)</li>
</ul>
</li>
<li class="">Run as:
<ul>
<li class="">FastAPI server (sync and async modes)</li>
<li class="">CLI (<code>llm-food convert-file my.pdf</code>)</li>
<li class="">Python client</li>
</ul>
</li>
<li class="">Async batch jobs with task tracking (DuckDB-based).</li>
<li class="">Docker-ready, with optional Bearer token auth.</li>
<li class="">Extensible and configurable.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="batch-mode-fast-cheap-parallel">Batch Mode: Fast, Cheap, Parallel<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#batch-mode-fast-cheap-parallel" class="hash-link" aria-label="Batch Mode: Fast, Cheap, Parallel doğrudan bağlantı" title="Batch Mode: Fast, Cheap, Parallel doğrudan bağlantı" translate="no">​</a></h3>
<p>PDFs get chunked and sent to Gemini's Batch API. You just upload files and give it a GCS path — we handle the queuing, status tracking, and Markdown export.</p>
<p>Other formats (DOCX, RTF, PPTX) are handled individually as background tasks.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-to-use-it">How to Use It<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#how-to-use-it" class="hash-link" aria-label="How to Use It doğrudan bağlantı" title="How to Use It doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="install">Install<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#install" class="hash-link" aria-label="Install doğrudan bağlantı" title="Install doğrudan bağlantı" translate="no">​</a></h3>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">pip </span><span class="token function" style="color:rgb(220, 220, 170)">install</span><span class="token plain"> </span><span class="token string" style="color:rgb(206, 145, 120)">'llm-food[server]'</span><span class="token plain">          </span><span class="token comment" style="color:rgb(106, 153, 85)"># Full server</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">pip </span><span class="token function" style="color:rgb(220, 220, 170)">install</span><span class="token plain"> </span><span class="token string" style="color:rgb(206, 145, 120)">'llm-food[server,pymupdf]'</span><span class="token plain">  </span><span class="token comment" style="color:rgb(106, 153, 85)"># Optional: if you want pymupdf backend</span><br></div></code></pre></div></div>
<p>Or just the client:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">uv </span><span class="token function" style="color:rgb(220, 220, 170)">add</span><span class="token plain"> llm-food</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="start-the-server">Start the Server<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#start-the-server" class="hash-link" aria-label="Start the Server doğrudan bağlantı" title="Start the Server doğrudan bağlantı" translate="no">​</a></h3>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">llm-food-serve</span><br></div></code></pre></div></div>
<p>API lives at <code>http://localhost:8000/docs</code>.</p>
<p>Want Docker?</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token function" style="color:rgb(220, 220, 170)">docker</span><span class="token plain"> build </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">-t</span><span class="token plain"> llm-food </span><span class="token builtin class-name" style="color:rgb(78, 201, 176)">.</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token function" style="color:rgb(220, 220, 170)">docker</span><span class="token plain"> run </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">-d</span><span class="token plain"> </span><span class="token parameter variable" style="color:rgb(156, 220, 254)">-p</span><span class="token plain"> </span><span class="token number" style="color:rgb(181, 206, 168)">8000</span><span class="token plain">:8000 --env-file .env llm-food</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="cli-usage">CLI Usage<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#cli-usage" class="hash-link" aria-label="CLI Usage doğrudan bağlantı" title="CLI Usage doğrudan bağlantı" translate="no">​</a></h3>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token plain">llm-food convert-file ./paper.pdf</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">llm-food convert-url https://example.com</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">llm-food batch-create doc1.pdf doc2.pdf gs://my-bucket/outputs/</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">llm-food batch-status </span><span class="token operator" style="color:rgb(212, 212, 212)">&lt;</span><span class="token plain">task_id</span><span class="token operator" style="color:rgb(212, 212, 212)">&gt;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain">llm-food batch-results </span><span class="token operator" style="color:rgb(212, 212, 212)">&lt;</span><span class="token plain">task_id</span><span class="token operator" style="color:rgb(212, 212, 212)">&gt;</span><br></div></code></pre></div></div>
<p>Set server URL/token via env:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#9CDCFE;--prism-background-color:#1E1E1E"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#9CDCFE;background-color:#1E1E1E"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#9CDCFE"><span class="token builtin class-name" style="color:rgb(78, 201, 176)">export</span><span class="token plain"> </span><span class="token assign-left variable" style="color:rgb(156, 220, 254)">LLMFOOD_SERVERURL</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">http://my-host:8000</span><br></div><div class="token-line" style="color:#9CDCFE"><span class="token plain"></span><span class="token builtin class-name" style="color:rgb(78, 201, 176)">export</span><span class="token plain"> </span><span class="token assign-left variable" style="color:rgb(156, 220, 254)">LLMFOOD_APITOKEN</span><span class="token operator" style="color:rgb(212, 212, 212)">=</span><span class="token plain">secret</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="auth--config">Auth &amp; Config<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#auth--config" class="hash-link" aria-label="Auth &amp; Config doğrudan bağlantı" title="Auth &amp; Config doğrudan bağlantı" translate="no">​</a></h3>
<p>Add a <code>.env</code> file to configure:</p>
<ul>
<li class="">GCS credentials</li>
<li class="">Gemini API location/project</li>
<li class="">File size limits</li>
<li class="">PDF backend</li>
<li class="">Auth token</li>
</ul>
<p>Run locally or on a containerized server.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="roadmap">Roadmap<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#roadmap" class="hash-link" aria-label="Roadmap doğrudan bağlantı" title="Roadmap doğrudan bağlantı" translate="no">​</a></h2>
<ul>
<li class="">Support for transcription of media files (videos and audios)</li>
<li class="">YouTube crawling and batch transcription</li>
<li class="">Configurable automatic chunking</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="final-thoughts">Final Thoughts<a href="https://altai.dev/tr/blog/llm-food-pdf-to-markdown#final-thoughts" class="hash-link" aria-label="Final Thoughts doğrudan bağlantı" title="Final Thoughts doğrudan bağlantı" translate="no">​</a></h2>
<p>Building a RAG pipeline, a search agent, or just need clean Markdown from messy enterprise docs? llm-food gets you there fast.</p>
<p>But let's be real: raw Markdown is just raw ingredients. We're cooking up an enterprise-grade engine that builds custom synthetic datasets and trains LLMs tailored to your docs. <a href="https://altai.dev/" target="_blank" rel="noopener noreferrer" class="">Join the waitlist</a> or drop us a line if that's what you're hungry for.</p>]]></content>
        <author>
            <name>Yusuf Sarıgöz</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Introducing AfterImage: Custom Vision Models Without the ML Overhead]]></title>
        <id>https://altai.dev/tr/blog/introducing-afterimage</id>
        <link href="https://altai.dev/tr/blog/introducing-afterimage"/>
        <updated>2025-04-01T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[Training a custom image classifier used to mean weeks of setup. AfterImage changes that — bring your data, get your model.]]></summary>
        <content type="html"><![CDATA[<p>Training a custom image classifier used to mean weeks of setup: picking a framework, writing data pipelines, managing GPU instances, tuning hyperparameters, and hoping the model actually generalizes.</p>
<p>AfterImage removes all of that.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-afterimage-does">What AfterImage does<a href="https://altai.dev/tr/blog/introducing-afterimage#what-afterimage-does" class="hash-link" aria-label="What AfterImage does doğrudan bağlantı" title="What AfterImage does doğrudan bağlantı" translate="no">​</a></h2>
<p>AfterImage is a no-code vision model trainer. You upload your images, label them, and click train. ALTAI handles the rest — model architecture selection, training infrastructure, optimization, and deployment.</p>
<p>The result: a production-ready REST API endpoint for your custom classifier.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="who-its-for">Who it's for<a href="https://altai.dev/tr/blog/introducing-afterimage#who-its-for" class="hash-link" aria-label="Who it's for doğrudan bağlantı" title="Who it's for doğrudan bağlantı" translate="no">​</a></h2>
<ul>
<li class=""><strong>Product teams</strong> who need visual inspection without hiring ML engineers</li>
<li class=""><strong>Healthcare providers</strong> training diagnostic image classifiers on proprietary datasets</li>
<li class=""><strong>Manufacturing companies</strong> building defect detection on their specific parts</li>
<li class=""><strong>Defense and security</strong> teams with sensitive data that can't leave their environment</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-it-works">How it works<a href="https://altai.dev/tr/blog/introducing-afterimage#how-it-works" class="hash-link" aria-label="How it works doğrudan bağlantı" title="How it works doğrudan bağlantı" translate="no">​</a></h2>
<ol>
<li class="">Upload your images (JPG/PNG, any resolution)</li>
<li class="">Label each image with a class</li>
<li class="">Click <strong>Start Training</strong></li>
<li class="">Get an API endpoint when training completes</li>
</ol>
<p>No Python. No CUDA drivers. No infrastructure management.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="try-it">Try it<a href="https://altai.dev/tr/blog/introducing-afterimage#try-it" class="hash-link" aria-label="Try it doğrudan bağlantı" title="Try it doğrudan bağlantı" translate="no">​</a></h2>
<p><a href="https://app.altai.dev/auth" target="_blank" rel="noopener noreferrer" class="">Create a free account</a> and train your first model today.</p>]]></content>
        <author>
            <name>ALTAI Team</name>
        </author>
    </entry>
    <entry>
        <title type="html"><![CDATA[Why Most Enterprise AI Projects Fail (And How to Fix It)]]></title>
        <id>https://altai.dev/tr/blog/why-enterprise-ai-fails</id>
        <link href="https://altai.dev/tr/blog/why-enterprise-ai-fails"/>
        <updated>2025-03-01T00:00:00.000Z</updated>
        <summary type="html"><![CDATA[According to Gartner, over 85% of AI projects never make it to production. After talking to hundreds of enterprise teams, the reasons are almost always the same.]]></summary>
        <content type="html"><![CDATA[<p>According to Gartner, over 85% of AI projects never make it to production. After talking to hundreds of enterprise teams, the reasons are almost always the same.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-three-failure-modes">The three failure modes<a href="https://altai.dev/tr/blog/why-enterprise-ai-fails#the-three-failure-modes" class="hash-link" aria-label="The three failure modes doğrudan bağlantı" title="The three failure modes doğrudan bağlantı" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-generic-models-on-specific-problems">1. Generic models on specific problems<a href="https://altai.dev/tr/blog/why-enterprise-ai-fails#1-generic-models-on-specific-problems" class="hash-link" aria-label="1. Generic models on specific problems doğrudan bağlantı" title="1. Generic models on specific problems doğrudan bağlantı" translate="no">​</a></h3>
<p>Most teams start with a foundation model (GPT-4, Gemini, Claude) and try to prompt-engineer their way to domain accuracy. It works for demos. It fails in production.</p>
<p>Your medical imaging data, your legal contracts, your manufacturing defects — these require models trained on <em>your</em> data, not on the internet.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-infrastructure-that-becomes-a-second-job">2. Infrastructure that becomes a second job<a href="https://altai.dev/tr/blog/why-enterprise-ai-fails#2-infrastructure-that-becomes-a-second-job" class="hash-link" aria-label="2. Infrastructure that becomes a second job doğrudan bağlantı" title="2. Infrastructure that becomes a second job doğrudan bağlantı" translate="no">​</a></h3>
<p>Building MLOps from scratch is a trap. Teams spend 80% of their time on infrastructure — managing training jobs, versioning models, scaling endpoints — instead of solving the actual business problem.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-no-path-from-prototype-to-production">3. No path from prototype to production<a href="https://altai.dev/tr/blog/why-enterprise-ai-fails#3-no-path-from-prototype-to-production" class="hash-link" aria-label="3. No path from prototype to production doğrudan bağlantı" title="3. No path from prototype to production doğrudan bağlantı" translate="no">​</a></h3>
<p>A notebook that works on your laptop is not a product. Most teams hit a wall when they try to scale from a POC to something their colleagues can actually use.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-works">What works<a href="https://altai.dev/tr/blog/why-enterprise-ai-fails#what-works" class="hash-link" aria-label="What works doğrudan bağlantı" title="What works doğrudan bağlantı" translate="no">​</a></h2>
<p>The enterprise AI projects that succeed share a pattern:</p>
<ul>
<li class=""><strong>Domain-specific data</strong> — they train on their own proprietary datasets</li>
<li class=""><strong>Managed infrastructure</strong> — they use platforms that abstract away the ops burden</li>
<li class=""><strong>API-first deployment</strong> — models are exposed as endpoints that existing systems can call</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-altai-addresses-this">How ALTAI addresses this<a href="https://altai.dev/tr/blog/why-enterprise-ai-fails#how-altai-addresses-this" class="hash-link" aria-label="How ALTAI addresses this doğrudan bağlantı" title="How ALTAI addresses this doğrudan bağlantı" translate="no">​</a></h2>
<p>ALTAI is built around this pattern. You bring your data. We handle training infrastructure, optimization, and deployment. The output is a production-ready API endpoint — not a notebook, not a prototype.</p>
<p><a href="https://altai.dev/platform" target="_blank" rel="noopener noreferrer" class="">See how it works →</a></p>]]></content>
        <author>
            <name>ALTAI Team</name>
        </author>
    </entry>
</feed>