Open Source

Turkish text processing

Akana

OPEN SOURCE

Build applications that account for the structure of Turkish.

Akana is a Turkish natural language processing toolkit built in Rust with Python bindings. Tokenization, morphology, normalization, readability, and text chunking provide a text-processing layer for search and language applications.

01 / INPUT TO OUTPUT
01 / Input

A Turkish paragraph containing inflected words, abbreviations, and multiple sentences.

02 / Process

Tokenize the text with Turkish-aware rules, then apply morphology, normalization, or chunking as needed.

03 / Output

Text segments and linguistic analyses for search, document review, or natural language processing workflows.

02 / WHAT IT DOES
01

Analyze Turkish word structure

Inspect morphological structure and handle Turkish-specific casing, suffixes, and spelling patterns.

02

Prepare text for retrieval

Split Turkish text into sentences or chunks while preserving their positions in the source text.

03

Measure and normalize text

Use Turkish-focused readability formulas, spelling checks, and tools for normalizing informal text.

03 / WORK WITH ALTAI

From a component
to your application.

ALTAI uses Akana to build the text-processing layer for Turkish document preparation and search workflows. We evaluate the relevant steps on your organization's text and integrate the selected operations into the application.

Explore custom development

WORK WITH ALTAI

Have a task for this?

Tell us what the system needs to do, what data is available, and where it should run.

Discuss this project