Kurdish, Arabic & Text Utilities
Kurdish Sorani and Arabic text handling done properly โ normalisation that respects the orthography, Iraqi phone and address parsing, transliteration, readability and PII redaction.
Text Encoding Cleanup
Repairs mojibake from double-decoded UTF-8, removes BOMs and zero-width characters, and normalises line endings, listing every fix applied.
Unicode Normalization
Applies NFC, NFD, NFKC or NFKD normalisation and reports the codepoints that changed.
Kurdish Sorani Normalizer
Normalises Kurdish Sorani to standard orthography โ Arabic yeh to Kurdish yeh, Arabic kaf to Kurdish kaf, correct heh/ae handling, Arabic-Indic to Latin numerals, tatweel and zero-width cleanup โ reporting every substitution.
Arabic Normalizer
Normalises Arabic text โ alef variants, teh marbuta, alef maqsura, diacritic removal, tatweel stripping and presentation-form folding.
Numeral System Converter
Converts between Latin, Arabic-Indic and Eastern Arabic-Indic digits in place, leaving surrounding text untouched.
Date Normalizer
Parses dates written in mixed Arabic, Kurdish and English forms โ including Arabic-Indic digits and localised month names โ into ISO-8601.
Iraqi Phone Normalizer
Normalises Iraqi phone numbers to E.164, identifying the carrier from the mobile prefix and validating length per network.
Iraqi Address Structuring
Parses free-text Iraqi and Kurdistan Region addresses into governorate, district, locality, street and landmark components with per-field confidence.
Currency String Parser
Parses currency amounts written with mixed symbols, separators, scale words and Arabic-Indic digits into a decimal value and ISO currency code.
Name Transliteration
Transliterates personal names between Arabic or Kurdish script and Latin, offering ranked variants for common ambiguous forms.
Keyword Extraction
Extracts keywords and phrases by TF-IDF and RAKE scoring with stopword lists for English, Arabic and Kurdish Sorani.
Language Detection
Detects language from script distribution and n-gram profiles, distinguishing Kurdish Sorani from Arabic and Persian, with per-candidate scores.
Lexicon Sentiment
Lexicon-based sentiment scoring with negation and intensifier handling, returning the matched terms so the score is auditable. No model inference involved.
Profanity Filter
Detects and masks profanity in English, Arabic and Kurdish including obfuscated spellings, with configurable masking.
PII Redaction
Detects and redacts emails, phone numbers, national IDs, card numbers with Luhn validation, IBANs and IP addresses, returning offsets and types but never the redacted values.
Text Diff
Produces line, word or character diffs with unified-diff output and a similarity ratio.
Readability Metrics
Computes Flesch Reading Ease, Flesch-Kincaid grade, Gunning Fog, SMOG, Coleman-Liau and ARI with the underlying syllable and sentence counts.
Text Statistics
Character, word, sentence and paragraph counts with script distribution, RTL share and top token frequencies.
Slug Generator
Generates URL-safe slugs from Latin, Arabic or Kurdish text with transliteration, length limits and collision suffixes.
Template Fill
Fills a placeholder template from a data object with strict unresolved-placeholder reporting. Sandboxed โ no expression evaluation or code execution.