Skip to content

Repository files navigation

Tunisian Arabic NLP Resources

Entries Open resources License PRs welcome DOI

A curated list of datasets, models, tools, and papers for the natural language processing of Tunisian Arabic (Tunisian Derja / Tounsi / تونسي, ISO 639-3 code aeb).

Also on Hugging Face: the inventory as a loadable table at datasets/fatmajlali/tunisian-nlp-resources (regenerated from these files by build-dataset-csv.py), and the open Tunisian datasets and models gathered in one place in the Tunisian Arabic (Derja) collection.

The goal is to be the single most complete inventory of what exists for Tunisian NLP: text and speech, open and gated, so that researchers, students, and engineers can find what is out there and see clearly where the gaps are.

This is meant to stay current, and that only works if it isn't maintained by one person. If you know a resource that is missing, or you built one, or something here is wrong or out of date:

Multi-dialect and pan-Arabic resources are welcome; the map records the Tunisian portion honestly rather than counting the whole thing.

At a glance

Entries Where
Text datasets, benchmarks, lexicons & papers 67 this file
Speech corpora (ASR, SLU, translation, TTS) 25 SPEECH.md
Pretrained models (LLMs, encoders, ASR, TTS) 16 MODELS.md
Researchers, labs & companies 30 PEOPLE.md

Counted as one ### heading each, so the figures above sum to the entries badge and anyone can reproduce them with grep -c '^### '. Two caveats in opposite directions: a few headings group several related items (for example "Classic ASR systems (papers)"), which undercounts individual resources; and a handful of resources are cross-listed under a second category with a pointer to the full entry (PADIC, TArC), which counts them twice. The figure is a heading count, not a claim about distinct artifacts.

126 access tags are applied across the three resource files: 96 open · 11 paywalled · 9 paper-only · 8 on request · 1 gated · 1 commercial. Some entries carry more than one tag (scripts open, underlying audio paywalled), so this counts tags rather than resources. Every entry links to a verifiable source; uncertain Tunisian coverage is flagged rather than dropped, and things checked and found to contain no Tunisian data are recorded under Confirmed negatives instead of silently omitted.

Recently added

Newest first. Contributed entries are credited to whoever pointed them out.

Date Entry Added by
2026-08 Whisperv3-tunisian-codeswitch — NADI 2026 subtask 1.3, TN↔FR/EN code-switched ASR; blind test WER 15.22 (3rd), against 0.478 for Tunisian on the 13-dialect model Ahmed Wasfy
2026-08 FARUKxAUTO/tunisian-asr-cleaned — 54k-row Tunisian ASR set behind the model above; flagged: no card, no licence, provenance unestablished Ahmed Wasfy
2026-08 tunisian-darija-english — 553 provenance-tagged Arabizi↔English pairs, 53 cultural categories + from-scratch MT pipeline Dhia Azizi
2026-07 dialect-router-v0.2 — 15-label Arabic dialect ID, 11.6M params; macro-F1 0.905, no per-dialect breakdown published Ahmed Wasfy
2026-07 whisper-large-v3-arabic-dialectal-v2 — 13-dialect ASR with per-dialect WER; Tunisian is hardest of the 13 (0.478) Ahmed Wasfy
2026-07 lahgtna-omnivoice-v2 — multi-dialect Arabic TTS, Tunisian supported Ahmed Wasfy
2026-07 lahgtna-v3-small — dialect-balanced ASR set, 200 Tunisian test clips Ahmed Wasfy
2026-07 Corrected CMN2 / NIST SRE — the ~396h figure covers Tunisian and English audio; Tunisian-only share unpublished maintainer

Contents

Separate files:

  • SPEECH.md — speech corpora (ASR, SLU, speech translation), TTS, speech dialect ID
  • MODELS.md — pretrained language models, ASR models, TTS models
  • PEOPLE.md — researchers, labs, companies

Access legend

Each entry ends with an access note:

  • [open] — freely downloadable, no login
  • [gated] — requires the owner's approval before access is granted
  • [on request] — obtain by contacting the authors
  • [paywalled] — behind a publisher or LDC paywall
  • [paper only] — described in a paper; no dataset download located
  • [commercial] — a paid product, not a research resource; listed for completeness

A note on scope: some resources are pan-Arabic or Maghrebi and contain Tunisian only as one part. These are included with the Tunisian portion noted, because they are often the only source of a given resource type. Items where Tunisian coverage could not be confirmed are flagged.


Text corpora (raw / web / social)

  • Raw reference corpus with a searchable concordance interface. Karen McNeil and Miled Faiza.
  • ~2,006 texts / ~882,000 words: folklore, songs, proverbs, blogs, email, Facebook, forums, transcribed radio.
  • [open] (online search interface; not a single bulk download).
  • Raw, unannotated Arabizi (Latin-script) corpus. Amara et al., 2021. Facebook public-page messages kept as-is, grouped by page/period. Companion to tunisiya.org.
  • [open] (Zenodo).
  • Aggregated raw corpus, ~802,659 text examples (~860k rows). Merges social media, conversational transcripts, chatbot dialogues, and other public Derja datasets; preserves French/English code-switching. Hamza Bouajila, 2025. Seed corpus for the ESPRIT-Derja model.
  • [open] (HuggingFace).
  • Large aggregated Derja LM corpus, ~2.23M rows / 332 MB, cc-by-sa-4.0. Wajdi Ghezaiel and Jean-Pierre Lorré (LINAGORA), 2025. Aggregates 14 sub-corpora (khaled123 Derja-English, TSAC, TunBERT, TunSwitch, TuDiCoI, QADI-TN, MADAR-TN, TA-Segmentation, Tweet_TN, and more). Used to continual-pretrain the Labess LLM.
  • [open] (HuggingFace).
  • The Tunisian (aeb) portion of FineWeb2, ~265k rows. LINAGORA, June 2025. ODC-By v1.0.
  • [open] (HuggingFace).
  • Raw social-media corpus scraped from Reddit. Small/uncurated; size not documented.
  • [open] (HuggingFace).

Sentiment analysis

  • Binary sentiment (positive/negative). Medhaffar, Bougares, Estève, Hadrich-Belguith, 2017.
  • ~17,000 Facebook comments from Tunisian radio/TV pages (2015–2016), mixed Arabic script and Arabizi. The most-cited Tunisian sentiment corpus.
  • Mirrors: HF fbougares/tsac, TensorFlow Datasets (tsac).
  • Paper: Sentiment Analysis of Tunisian Dialects (WANLP 2017). [open].
  • Extended Arabizi sentiment set, ~100k comments (movies, politics, sport…) labeled positive/negative/neutral. iCompass, 2021.
  • Distributed via a Zindi challenge and on Kaggle. No single canonical repo for the full 100k.
  • [open].
  • Polarity classification (negative/positive), 49,889 rows, Arabic + Arabizi. Aggregated Tunisian tweets/comments.
  • [open] (HuggingFace).
  • Sentiment corpus, ~23,786 rows. Hedi Naouara (LINAGORA), Oct 2024.
  • [open] (HuggingFace).

Related: learning word representations for Tunisian sentiment (arXiv:2010.06857).


Offensive language, hate speech, sarcasm

  • Three classes: normal / abusive / hate. Mulki, Haddad et al., 2019. ~6,075 comments.
  • Paper: Springer chapter. [open] (GitHub).
  • Toxic-speech corpus, ~10,000 comments. Gharbi, Haddad, Kchaou, Arfaoui, 2021. Paper is open; a public data repo was not confirmed.
  • [paper only] (arXiv:2110.05287).

HateTune — Tunisian Dialect Hate Speech Detection Dataset

  • Hate speech in Arabic-script Tunisian, 2024. Behind a Springer paywall; dataset access unconfirmed.
  • [paywalled] (Springer chapter).

Dialect identification (text)

Datasets and benchmarks. Trained dialect-ID models are in MODELS.md § Dialect identification.

  • Multi-dialect parallel corpus + city/country dialect ID (26-way). Bouamor et al., LREC 2018. Includes Tunis (Corpus-26/Corpus-6) and Sfax in the lexicon.
  • Paper: LREC 2018. [open] (free research license via form).
  • Country/province-level dialect ID from tweets, annual 2020–2024. Tunisia is one of ~21 covered countries.
  • [open] (via shared-task registration). See Shared tasks.
  • Aggregated dialect-ID dataset merging DART, SHAMI, PADIC, AOC, and TSAC — so it embeds a Tunisian subset. Zahir et al., 2021.
  • Data paper: ScienceDirect. [open] (GitHub).
  • Country-level dialect-ID tweets across 18 countries; the Tunisian subset is ~8,879 items (as folded into the LinTO Derja aggregate). QCRI.
  • [open] (QCRI resources page).
  • Large multi-dialect Arabic Twitter corpus for gender, age, and language-variety identification, covering 11 Arab regions including Tunisia. Zaghouani and Charfi, LREC 2018.
  • [on request] (contact authors); Tunisian is one regional subset.

Multi-Dialect, Multi-Genre Corpus of Informal Written Arabic (LREC 2014)

  • Cotterell and Callison-Burch. Informal Arabic across five dialect groups including Maghrebi (Tunisian-relevant). Paper: LREC 2014. [on request].
  • Dialect-ID train/test data including Maghrebi/Tunisian-relevant material.
  • [open] (GitHub).

Sub-dialect identification (within Tunisian)


Treebanks and syntactic resources

  • Multi-layer annotated Arabizi corpus. Gugliotta and Dinarelli; first release 2020, complete release 2022.
  • 4,797 sentences / ~43,300 tokens (forum, social, blog, rap), with user metadata (governorate, age range, gender).
  • Annotation layers: token classification (arabizi/foreign/emotag), Arabic-script transliteration (CODA-TUN), tokenization with clitic splitting, POS (Penn Arabic Treebank style), lemmatization (in progress). CC BY-NC-SA 4.0.
  • Companion annotation tool: Multi-Task Sequence Prediction System (GitLab).
  • Papers: LREC 2022 (arXiv:2207.04796), LREC 2020, WANLP 2020. Mirror: HF arbml/TArC. [open].
  • Constituency/syntactic treebank and a trained Stanford-parser model. Mekki, Zribi, Ellouze, Belguith (ANLP-RG, Sfax), 2020.
  • Tunisian constitution (12,378 words) + 1,072 STAC sentences, CODA-TUN orthography; parser trained on ~8,000 sentences / ~79,600 tokens, best F-measure 80.12%.
  • Paper: AICCSA 2020. [open] (GitHub; IEEE paper paywalled).

TADT — Tunisian Arabic Dependency Treebank (Universal Dependencies)

  • The first UD treebank for Tunisian (aeb). Amal Aissaoui (CUNY), 2026. 100 sentences / 1,466 tokens of Arabizi social-media text (sampled from CTAB), with UPOS, morphological features, lemmas, and dependency relations. Built by cross-dialect transfer from the Algerian NArabizi treebank + manual correction.
  • The treebank is being finalized for official UD release; not yet downloadable.
  • [paper only] (UDW 2026 paper).
  • Sentence-boundary / segmentation annotation with CODA-TA normalization. Mekki et al., 2021. 260,364 words / 33,581 sentences, expert-validated.
  • [open] (GitHub).

POS tagging

TArC POS layer

  • See TArC (repo): ~43k Arabizi tokens with Penn Arabic Treebank-style POS. [open].
  • Hamdi, Nasr, Habash, Gala. Maps Tunisian to an MSA lattice and tags with an MSA tagger (~89% accuracy). Method paper; no standalone gold TD POS corpus released. [open] (paper).

Fine-Grained POS Tagging of Spoken Tunisian Dialect (Springer NLDB 2014)

  • Boujelbane, Mallek, Ellouze, Hadrich-Belguith. Builds a TD training corpus by converting an MSA corpus via a bilingual lexicon (~78.5% accuracy).
  • [paywalled] (Springer).

Morphology (analyzers and disambiguation)

  • 1,000 Tunisian words used to evaluate the Karmani analyzer. Nadia Ben Mohamed Karmani (REGIM, Sfax), ~2016.
  • Paper: AICCSA 2016 (IEEE). [open] (GitHub; paper paywalled).
  • XML lexicon of the main Tunisian clitics, used by the Karmani analyzer/tokenizer.
  • [open] (GitHub).

Al-Khalil-TUN — Morphological Analysis of Tunisian Dialect (IJCNLP 2013)

  • Zribi, Ellouze Khemakhem, Hadrich Belguith. Adapts the MSA Al-Khalil analyzer with a purpose-built Tunisian lexicon (~30k words). Analyzer and lexicon not publicly downloadable.
  • [paper only] (paper).
  • Zribi, Ellouze, Hadrich-Belguith, Blache. ML disambiguation over the analyzer output (~87–88% accuracy). Journal of King Saud University – Computer and Information Sciences 29(2), pp. 147–155. [open] (article).

Named Entity Recognition

The most useful openly available NER building blocks are the Barcha gazetteers (below). A token-annotated gold NER corpus specific to Tunisian has not been openly released yet. Known work:

  • Open multi-purpose Tunisian resource (MIT). wa3dbk. Its named_entities/ folder is a set of curated Tunisian entity gazetteers: people (academics/scientists, artists, footballers, media figures, poets, politicians, trade unionists, writers), institutions/associations/companies, government institutions, cities, universities (public and private), ISETs, political parties, and unions. Also ships a texts/ raw-text folder and a translation/ folder (Tunisian↔English/MSA).
  • These are gazetteers/word lists rather than a token-annotated corpus, but they are the best open NER starting point for Tunisian. [open].

TUNER — NER of Tunisian Arabic with Bi-LSTM-CRF (2023)

  • Hybrid Bi-LSTM-CRF + rules, F-measure 91.43%; involves an annotated Tunisian NER corpus that is not publicly released.
  • [paywalled] (World Scientific).

Recognition and Translation of Tunisian Dialect Named Entities into MSA (2021)

  • Torjmen and Haddar. No public dataset found.
  • [paper only] (ResearchGate).

Machine translation and parallel corpora

  • Parallel dialect corpus. Meftouh, Harrat, Abbas, Smaïli (SMarT/LORIA); Tunisian portion by Salma Jamoussi. 2015, extended 2017–2018.
  • ~6,400 sentences per variety aligned to MSA. Varieties: Algiers, Annaba, Tunisian, Moroccan (Casablanca, Rabat), Syrian, Palestinian + MSA.
  • Direct download: ZIP; SourceForge mirror.
  • Papers: PACLIC 2015, ICAT 2018. [open].
  • Corpus-26: 2,000 BTEC sentences translated into 25 city dialects + MSA (includes Tunis and Sfax). Corpus-6: 12,000 sentences in 5 cities + MSA (includes Tunis). Bouamor et al., LREC 2018.
  • [open] (free research license via form).
  • 2,000 sentences in MSA, Egyptian, Tunisian, Jordanian, Palestinian, Syrian + English. Bouamor, Habash, Oflazer, LREC 2014.
  • [on request] (paper open; no confirmed open direct download).
  • Kchaou, Boujelbane, Hadrich-Belguith. Tunisian (social media) ↔ MSA with data augmentation; BLEU up to 15.03.
  • [paper only] (dataset link not clearly published).

Rule-based / SMT Tunisian→MSA

  • 553 Tunisian Arabizi↔English pairs across 53 cultural categories (louage culture, mawsem el zitoun, bac exam culture, el 3aza w mawt…). Every pair provenance-tagged in a source column: 500 self-written by the author (native speaker), 53 field-collected from family and community speakers with documented consent; nothing synthetic. Dhia Azizi, 2026; collection ongoing.
  • Companion repo: darija-translator — a from-scratch ~15.6M-param encoder-decoder with an Arabizi-aware BPE tokenizer (3/7/9/5 as protected markers), pretrained on cleaned Moroccan darija (atlasia/darija_english) and fine-tuned on the Tunisian pairs. Locked test set; BLEU 3.89 reported honestly as the v1 baseline. The Tunisian data is the 553 pairs — the ~36k pretraining pairs are Moroccan-derived.
  • [open] (CC BY-NC-SA 4.0 — note the non-commercial clause).

Community HuggingFace parallel sets

Speech translation corpora with Tunisian↔English/French text

  • TuniFra (Tunisian↔French, 15h), TEDxTN and IWSLT 2022/2023 Tunisian–English are three-way (audio + transcript + translation) and are listed in SPEECH.md; their transcript/translation layers are also Tunisian–French / Tunisian–English bitext.

Transliteration and Arabizi

  • Annotated Arabizi corpus and the neural tool that transliterates Arabizi→CODA Arabic script and does tokenization/POS. Gugliotta, Dinarelli, Kraif. Code on GitLab. [open].

Masmoudi et al. — Arabizi→Arabic-script transliteration for Tunisian

Younes et al. — bi-script Tunisian resources


Orthography (CODA)

  • The canonical Tunisian CODA spec (Arabic-script conventional orthography). Zribi, Boujelbane, Masmoudi, Ellouze, Belguith, Habash, LREC 2014. PDF. [open].
  • Cross-dialect orthography framework (covers Tunisian). Habash et al., LREC 2018. PDF. [open].
  • Turki, Adel, Gibson, Zribi. Extends CODA across Maghrebi (incl. Tunisian). [open].

NOTA — Normalized Orthography for Tunisian Arabic (Springer 2025)

  • Adaptation/normalization of CODA* for Tunisian. [paywalled] (Springer).

Lexicons, dictionaries, wordnets

  • Tunisian (aeb) WordNet in ISO-LMF format. Karmani, Soussou, Alimi (REGIM, Sfax), 2015. Built from the Peace Corps English–TA dictionary and projected from Princeton WordNet 3.1 (~18,209 synsets). [open] (GitHub; ACLing 2015 paper paywalled).

TunDiaWN — Tunisian dialect WordNet (WANLP 2014)

  • Bouchlaghem and Elkhlifi. Corpus-based wordnet reusing English + Arabic wordnets. Database not found for download.
  • [paper only] (WANLP 2014).
  • 1977 two-way English↔Tunisian dictionary with phonetics and grammar notes; the seed lexicon (5,133 words) for aebWordNet. Companion Peace Corps Tunisian Arabic course. [open] (Internet Archive).
  • Online Tunisian Arabic dictionary. [open] (web).

Bilingual lexicons from the literature (Tunisian↔MSA)


LLM evaluation benchmarks

  • The dedicated Tunisian LLM instruction benchmark. Ben Hassine, Arrak, Addhoum, Wilson (CoCoA Lab, Univ. of Michigan-Flint), EMNLP 2025.
  • 744 Tunisian instructions with gold human responses; ten Arabic-claiming LLMs scored by humans and an LLM judge on quality, correctness, relevance, dialectal adherence. Finding: most LLMs struggle in Tunisian.
  • Repo: Souha-BH/TounsiBench-… (data/ = instructions, native gold responses, topic labels, evaluated model outputs; plus a GPT-4o evaluation notebook to score your own model). Paper: EMNLP 2025. [open].
  • Parallel Tunizi ↔ standard Tunisian ↔ English corpus with sentiment labels; benchmarks LLMs on transliteration, translation, sentiment. Mahdi et al., 2025. [open] (arXiv).
  • MMLU-style multiple-choice evaluation benchmark in Tunisian Derja, ~21.7k rows. LINAGORA, 2025. Built to evaluate the Labess LLM in-dialect.
  • [open] (HuggingFace).
  • Benchmark dataset for Tunisian/local dialects, LLM-eval oriented. Size/tasks not fully verified. [open] (HuggingFace).

LLM training and evaluation datasets

The instruction / SFT / DPO / synthetic data layer behind the Tunisian LLMs in MODELS.md. Mostly 2025–2026, mostly open, and mostly unreviewed community data, quality varies, so treat sizes as raw counts.

LINAGORA / Labess training stack

ESPRIT

  • ESPRIT-Derja-Instruct — ~7,013 bilingual Arabic/Arabizi Tunisian instruction pairs distilled from GPT-4o; trains the ESPRIT-Derja-Qwen3-8B model. [open] (bundled with the model repo).

Tunisia.AI / khaled123 (Khaled Bouzaiene)

Other


Surveys


Shared tasks


Other resource lists


How to cite

If this inventory is useful in your research, please cite it:

@misc{jlali2026tunisiannlp,
  author       = {Jlali, Fatma},
  title        = {Tunisian Arabic {NLP} Resources: A Curated Inventory},
  year         = {2026},
  howpublished = {\url{https://github.com/jjlalli/Tunisian-NLP-Resources}},
  note         = {Living inventory of datasets, models, and papers for Tunisian Arabic (aeb)}
}

A paper describing this inventory is in preparation; this entry will be updated when it is available.

License

The list itself is released under CC BY 4.0. The linked resources keep their own licenses — check each entry.


Maintained by Fatma Jlali. Contributions and corrections welcome — see CONTRIBUTING.md. Last compiled: July 2026.

About

Open, maintained inventory of NLP resources for Tunisian Arabic (aeb). Contributions welcome.

Topics

Resources

Contributing

Stars

20 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages