A curated list of datasets, models, tools, and papers for the natural language processing of Tunisian Arabic (Tunisian Derja / Tounsi / تونسي, ISO 639-3 code aeb).
Also on Hugging Face: the inventory as a loadable table at datasets/fatmajlali/tunisian-nlp-resources (regenerated from these files by build-dataset-csv.py), and the open Tunisian datasets and models gathered in one place in the Tunisian Arabic (Derja) collection.
The goal is to be the single most complete inventory of what exists for Tunisian NLP: text and speech, open and gated, so that researchers, students, and engineers can find what is out there and see clearly where the gaps are.
This is meant to stay current, and that only works if it isn't maintained by one person. If you know a resource that is missing, or you built one, or something here is wrong or out of date:
- Add a resource — paste a link, that's enough; formatting and verification are my job
- Report something wrong — dead links and overstated numbers make this worse than useless
- Or open a pull request directly — see CONTRIBUTING.md
Multi-dialect and pan-Arabic resources are welcome; the map records the Tunisian portion honestly rather than counting the whole thing.
| Entries | Where | |
|---|---|---|
| Text datasets, benchmarks, lexicons & papers | 67 | this file |
| Speech corpora (ASR, SLU, translation, TTS) | 25 | SPEECH.md |
| Pretrained models (LLMs, encoders, ASR, TTS) | 16 | MODELS.md |
| Researchers, labs & companies | 30 | PEOPLE.md |
Counted as one ### heading each, so the figures above sum to the entries badge and anyone can reproduce them with grep -c '^### '. Two caveats in opposite directions: a few headings group several related items (for example "Classic ASR systems (papers)"), which undercounts individual resources; and a handful of resources are cross-listed under a second category with a pointer to the full entry (PADIC, TArC), which counts them twice. The figure is a heading count, not a claim about distinct artifacts.
126 access tags are applied across the three resource files: 96 open · 11 paywalled · 9 paper-only · 8 on request · 1 gated · 1 commercial. Some entries carry more than one tag (scripts open, underlying audio paywalled), so this counts tags rather than resources. Every entry links to a verifiable source; uncertain Tunisian coverage is flagged rather than dropped, and things checked and found to contain no Tunisian data are recorded under Confirmed negatives instead of silently omitted.
Newest first. Contributed entries are credited to whoever pointed them out.
| Date | Entry | Added by |
|---|---|---|
| 2026-08 | Whisperv3-tunisian-codeswitch — NADI 2026 subtask 1.3, TN↔FR/EN code-switched ASR; blind test WER 15.22 (3rd), against 0.478 for Tunisian on the 13-dialect model | Ahmed Wasfy |
| 2026-08 | FARUKxAUTO/tunisian-asr-cleaned — 54k-row Tunisian ASR set behind the model above; flagged: no card, no licence, provenance unestablished | Ahmed Wasfy |
| 2026-08 | tunisian-darija-english — 553 provenance-tagged Arabizi↔English pairs, 53 cultural categories + from-scratch MT pipeline | Dhia Azizi |
| 2026-07 | dialect-router-v0.2 — 15-label Arabic dialect ID, 11.6M params; macro-F1 0.905, no per-dialect breakdown published | Ahmed Wasfy |
| 2026-07 | whisper-large-v3-arabic-dialectal-v2 — 13-dialect ASR with per-dialect WER; Tunisian is hardest of the 13 (0.478) | Ahmed Wasfy |
| 2026-07 | lahgtna-omnivoice-v2 — multi-dialect Arabic TTS, Tunisian supported | Ahmed Wasfy |
| 2026-07 | lahgtna-v3-small — dialect-balanced ASR set, 200 Tunisian test clips | Ahmed Wasfy |
| 2026-07 | Corrected CMN2 / NIST SRE — the ~396h figure covers Tunisian and English audio; Tunisian-only share unpublished | maintainer |
- Text corpora (raw / web / social)
- Sentiment analysis
- Offensive language, hate speech, sarcasm
- Dialect identification (text)
- Treebanks and syntactic resources
- POS tagging
- Morphology (analyzers and disambiguation)
- Named Entity Recognition
- Machine translation and parallel corpora
- Transliteration and Arabizi
- Orthography (CODA)
- Lexicons, dictionaries, wordnets
- LLM evaluation benchmarks
- LLM training and evaluation datasets
- Surveys
- Shared tasks
- Other resource lists
Separate files:
- SPEECH.md — speech corpora (ASR, SLU, speech translation), TTS, speech dialect ID
- MODELS.md — pretrained language models, ASR models, TTS models
- PEOPLE.md — researchers, labs, companies
Each entry ends with an access note:
- [open] — freely downloadable, no login
- [gated] — requires the owner's approval before access is granted
- [on request] — obtain by contacting the authors
- [paywalled] — behind a publisher or LDC paywall
- [paper only] — described in a paper; no dataset download located
- [commercial] — a paid product, not a research resource; listed for completeness
A note on scope: some resources are pan-Arabic or Maghrebi and contain Tunisian only as one part. These are included with the Tunisian portion noted, because they are often the only source of a given resource type. Items where Tunisian coverage could not be confirmed are flagged.
- Raw reference corpus with a searchable concordance interface. Karen McNeil and Miled Faiza.
- ~2,006 texts / ~882,000 words: folklore, songs, proverbs, blogs, email, Facebook, forums, transcribed radio.
- [open] (online search interface; not a single bulk download).
- Raw, unannotated Arabizi (Latin-script) corpus. Amara et al., 2021. Facebook public-page messages kept as-is, grouped by page/period. Companion to tunisiya.org.
- [open] (Zenodo).
- Aggregated raw corpus, ~802,659 text examples (~860k rows). Merges social media, conversational transcripts, chatbot dialogues, and other public Derja datasets; preserves French/English code-switching. Hamza Bouajila, 2025. Seed corpus for the ESPRIT-Derja model.
- [open] (HuggingFace).
- Large aggregated Derja LM corpus, ~2.23M rows / 332 MB, cc-by-sa-4.0. Wajdi Ghezaiel and Jean-Pierre Lorré (LINAGORA), 2025. Aggregates 14 sub-corpora (khaled123 Derja-English, TSAC, TunBERT, TunSwitch, TuDiCoI, QADI-TN, MADAR-TN, TA-Segmentation, Tweet_TN, and more). Used to continual-pretrain the Labess LLM.
- [open] (HuggingFace).
- The Tunisian (
aeb) portion of FineWeb2, ~265k rows. LINAGORA, June 2025. ODC-By v1.0. - [open] (HuggingFace).
- Raw social-media corpus scraped from Reddit. Small/uncurated; size not documented.
- [open] (HuggingFace).
- Binary sentiment (positive/negative). Medhaffar, Bougares, Estève, Hadrich-Belguith, 2017.
- ~17,000 Facebook comments from Tunisian radio/TV pages (2015–2016), mixed Arabic script and Arabizi. The most-cited Tunisian sentiment corpus.
- Mirrors: HF fbougares/tsac, TensorFlow Datasets (
tsac). - Paper: Sentiment Analysis of Tunisian Dialects (WANLP 2017). [open].
- Sentiment on Latin-script (Arabizi) Tunisian. Fourati, Messaoudi, Haddad (iCompass), 2020. ~9k YouTube comments.
- Mirrors: Zenodo 4275240, iCompass-ai/TUNIZI, HF chaymafourati/tunizi, arbml/TUNIZI.
- Paper: arXiv:2004.14303. [open].
- Extended Arabizi sentiment set, ~100k comments (movies, politics, sport…) labeled positive/negative/neutral. iCompass, 2021.
- Distributed via a Zindi challenge and on Kaggle. No single canonical repo for the full 100k.
- [open].
- Polarity classification (negative/positive), 49,889 rows, Arabic + Arabizi. Aggregated Tunisian tweets/comments.
- [open] (HuggingFace).
- Sentiment corpus, ~23,786 rows. Hedi Naouara (LINAGORA), Oct 2024.
- [open] (HuggingFace).
Related: learning word representations for Tunisian sentiment (arXiv:2010.06857).
- Three classes: normal / abusive / hate. Mulki, Haddad et al., 2019. ~6,075 comments.
- Paper: Springer chapter. [open] (GitHub).
- Toxic-speech corpus, ~10,000 comments. Gharbi, Haddad, Kchaou, Arfaoui, 2021. Paper is open; a public data repo was not confirmed.
- [paper only] (arXiv:2110.05287).
- Hate speech in Arabic-script Tunisian, 2024. Behind a Springer paywall; dataset access unconfirmed.
- [paywalled] (Springer chapter).
Datasets and benchmarks. Trained dialect-ID models are in MODELS.md § Dialect identification.
- Multi-dialect parallel corpus + city/country dialect ID (26-way). Bouamor et al., LREC 2018. Includes Tunis (Corpus-26/Corpus-6) and Sfax in the lexicon.
- Paper: LREC 2018. [open] (free research license via form).
- Country/province-level dialect ID from tweets, annual 2020–2024. Tunisia is one of ~21 covered countries.
- [open] (via shared-task registration). See Shared tasks.
- Aggregated dialect-ID dataset merging DART, SHAMI, PADIC, AOC, and TSAC — so it embeds a Tunisian subset. Zahir et al., 2021.
- Data paper: ScienceDirect. [open] (GitHub).
- Country-level dialect-ID tweets across 18 countries; the Tunisian subset is ~8,879 items (as folded into the LinTO Derja aggregate). QCRI.
- [open] (QCRI resources page).
- Large multi-dialect Arabic Twitter corpus for gender, age, and language-variety identification, covering 11 Arab regions including Tunisia. Zaghouani and Charfi, LREC 2018.
- [on request] (contact authors); Tunisian is one regional subset.
- Cotterell and Callison-Burch. Informal Arabic across five dialect groups including Maghrebi (Tunisian-relevant). Paper: LREC 2014. [on request].
- Dialect-ID train/test data including Maghrebi/Tunisian-relevant material.
- [open] (GitHub).
- Cross-listed. Also usable for dialect ID; one of six dialects is Tunisian. Full details under Machine translation and parallel corpora. [open].
- Text and Speech-based Tunisian Arabic Sub-Dialects Identification (LREC 2020) — Kchaou, Ben Abdallah, Bougares. Distinguishes Tunis / Sfax / Sousse / Tataouine. A released benchmark rather than an organized competition. [open] (paper).
- Multi-layer annotated Arabizi corpus. Gugliotta and Dinarelli; first release 2020, complete release 2022.
- 4,797 sentences / ~43,300 tokens (forum, social, blog, rap), with user metadata (governorate, age range, gender).
- Annotation layers: token classification (arabizi/foreign/emotag), Arabic-script transliteration (CODA-TUN), tokenization with clitic splitting, POS (Penn Arabic Treebank style), lemmatization (in progress). CC BY-NC-SA 4.0.
- Companion annotation tool: Multi-Task Sequence Prediction System (GitLab).
- Papers: LREC 2022 (arXiv:2207.04796), LREC 2020, WANLP 2020. Mirror: HF arbml/TArC. [open].
- Constituency/syntactic treebank and a trained Stanford-parser model. Mekki, Zribi, Ellouze, Belguith (ANLP-RG, Sfax), 2020.
- Tunisian constitution (12,378 words) + 1,072 STAC sentences, CODA-TUN orthography; parser trained on ~8,000 sentences / ~79,600 tokens, best F-measure 80.12%.
- Paper: AICCSA 2020. [open] (GitHub; IEEE paper paywalled).
- The first UD treebank for Tunisian (
aeb). Amal Aissaoui (CUNY), 2026. 100 sentences / 1,466 tokens of Arabizi social-media text (sampled from CTAB), with UPOS, morphological features, lemmas, and dependency relations. Built by cross-dialect transfer from the Algerian NArabizi treebank + manual correction. - The treebank is being finalized for official UD release; not yet downloadable.
- [paper only] (UDW 2026 paper).
- Sentence-boundary / segmentation annotation with CODA-TA normalization. Mekki et al., 2021. 260,364 words / 33,581 sentences, expert-validated.
- [open] (GitHub).
- Hamdi, Nasr, Habash, Gala. Maps Tunisian to an MSA lattice and tags with an MSA tagger (~89% accuracy). Method paper; no standalone gold TD POS corpus released. [open] (paper).
- Boujelbane, Mallek, Ellouze, Hadrich-Belguith. Builds a TD training corpus by converting an MSA corpus via a bilingual lexicon (~78.5% accuracy).
- [paywalled] (Springer).
- 1,000 Tunisian words used to evaluate the Karmani analyzer. Nadia Ben Mohamed Karmani (REGIM, Sfax), ~2016.
- Paper: AICCSA 2016 (IEEE). [open] (GitHub; paper paywalled).
- XML lexicon of the main Tunisian clitics, used by the Karmani analyzer/tokenizer.
- [open] (GitHub).
- Zribi, Ellouze Khemakhem, Hadrich Belguith. Adapts the MSA Al-Khalil analyzer with a purpose-built Tunisian lexicon (~30k words). Analyzer and lexicon not publicly downloadable.
- [paper only] (paper).
- Zribi, Ellouze, Hadrich-Belguith, Blache. ML disambiguation over the analyzer output (~87–88% accuracy). Journal of King Saud University – Computer and Information Sciences 29(2), pp. 147–155. [open] (article).
The most useful openly available NER building blocks are the Barcha gazetteers (below). A token-annotated gold NER corpus specific to Tunisian has not been openly released yet. Known work:
- Open multi-purpose Tunisian resource (MIT). wa3dbk. Its
named_entities/folder is a set of curated Tunisian entity gazetteers: people (academics/scientists, artists, footballers, media figures, poets, politicians, trade unionists, writers), institutions/associations/companies, government institutions, cities, universities (public and private), ISETs, political parties, and unions. Also ships atexts/raw-text folder and atranslation/folder (Tunisian↔English/MSA). - These are gazetteers/word lists rather than a token-annotated corpus, but they are the best open NER starting point for Tunisian. [open].
- Hybrid Bi-LSTM-CRF + rules, F-measure 91.43%; involves an annotated Tunisian NER corpus that is not publicly released.
- [paywalled] (World Scientific).
- Torjmen and Haddar. No public dataset found.
- [paper only] (ResearchGate).
- Parallel dialect corpus. Meftouh, Harrat, Abbas, Smaïli (SMarT/LORIA); Tunisian portion by Salma Jamoussi. 2015, extended 2017–2018.
- ~6,400 sentences per variety aligned to MSA. Varieties: Algiers, Annaba, Tunisian, Moroccan (Casablanca, Rabat), Syrian, Palestinian + MSA.
- Direct download: ZIP; SourceForge mirror.
- Papers: PACLIC 2015, ICAT 2018. [open].
- Corpus-26: 2,000 BTEC sentences translated into 25 city dialects + MSA (includes Tunis and Sfax). Corpus-6: 12,000 sentences in 5 cities + MSA (includes Tunis). Bouamor et al., LREC 2018.
- [open] (free research license via form).
- 2,000 sentences in MSA, Egyptian, Tunisian, Jordanian, Palestinian, Syrian + English. Bouamor, Habash, Oflazer, LREC 2014.
- [on request] (paper open; no confirmed open direct download).
- Kchaou, Boujelbane, Hadrich-Belguith. Tunisian (social media) ↔ MSA with data augmentation; BLEU up to 15.03.
- [paper only] (dataset link not clearly published).
- Rule-Based MT from Tunisian to MSA (Procedia 2020) — Sghaier and Zrigui. [open].
- FST + seq2seq Transformer Tunisian→MSA (ACM TALLIP 2024) — BLEU 56.65 (FST) / 66.07 (transformer). [paywalled].
- Phrase-based SMT + 5,000-sentence Tunis↔MSA corpus (IBIMA) — Sghaier and Zrigui. [on request].
- 553 Tunisian Arabizi↔English pairs across 53 cultural categories (louage culture, mawsem el zitoun, bac exam culture, el 3aza w mawt…). Every pair provenance-tagged in a
sourcecolumn: 500 self-written by the author (native speaker), 53 field-collected from family and community speakers with documented consent; nothing synthetic. Dhia Azizi, 2026; collection ongoing. - Companion repo: darija-translator — a from-scratch ~15.6M-param encoder-decoder with an Arabizi-aware BPE tokenizer (3/7/9/5 as protected markers), pretrained on cleaned Moroccan darija (atlasia/darija_english) and fine-tuned on the Tunisian pairs. Locked test set; BLEU 3.89 reported honestly as the v1 baseline. The Tunisian data is the 553 pairs — the ~36k pretraining pairs are Moroccan-derived.
- [open] (CC BY-NC-SA 4.0 — note the non-commercial clause).
- tunis-ai/tunisian-msa-parallel-corpus and tunis-ai/MADAR-TUN (~30k) — Derja↔MSA. Tunisia.AI, 2025. [open].
- NadiaGHEZAIEL/English_to_Tunisian_Dataset (~1.7k) — English→Tunisian. Nov 2025. [open].
- khaled123/Tunisianderjasynthtranslation (~436k, synthetic) and the Kaggle drejja-to-english (~13k) — Derja↔English. [open].
- Barcha — its
translation/folder holds Tunisian↔English/MSA parallel data (also has NER gazetteers and raw texts; see NER). [open].
- TuniFra (Tunisian↔French, 15h), TEDxTN and IWSLT 2022/2023 Tunisian–English are three-way (audio + transcript + translation) and are listed in SPEECH.md; their transcript/translation layers are also Tunisian–French / Tunisian–English bitext.
- Annotated Arabizi corpus and the neural tool that transliterates Arabizi→CODA Arabic script and does tokenization/POS. Gugliotta, Dinarelli, Kraif. Code on GitLab. [open].
- Transliteration of Arabizi into Arabic Script for Tunisian Dialect (ACM TALLIP 2019) — rule-based + CRF, WER 14.55%. [paywalled].
- Preliminary investigation (CICLing/Springer 2015) and an HMM approach. [paywalled].
- Romanized Tunisian Dialect Transliteration using Sequence Labelling (J. King Saud Univ. 2020). [open] (paper).
- Building Bi-script Language Resources for the Tunisian Dialect (Procedia 2021) — 284,894 Romanized messages; two bi-script dictionaries (Latin→Arabic 293,570 entries; Arabic→Latin 155,954 entries). Data not openly hosted. [paper only].
- Seq2Seq double transliteration (Procedia 2018). [open] (paper).
- The canonical Tunisian CODA spec (Arabic-script conventional orthography). Zribi, Boujelbane, Masmoudi, Ellouze, Belguith, Habash, LREC 2014. PDF. [open].
- Cross-dialect orthography framework (covers Tunisian). Habash et al., LREC 2018. PDF. [open].
- Turki, Adel, Gibson, Zribi. Extends CODA across Maghrebi (incl. Tunisian). [open].
- Adaptation/normalization of CODA* for Tunisian. [paywalled] (Springer).
- Tunisian (
aeb) WordNet in ISO-LMF format. Karmani, Soussou, Alimi (REGIM, Sfax), 2015. Built from the Peace Corps English–TA dictionary and projected from Princeton WordNet 3.1 (~18,209 synsets). [open] (GitHub; ACLing 2015 paper paywalled).
- Bouchlaghem and Elkhlifi. Corpus-based wordnet reusing English + Arabic wordnets. Database not found for download.
- [paper only] (WANLP 2014).
- 1977 two-way English↔Tunisian dictionary with phonetics and grammar notes; the seed lexicon (5,133 words) for aebWordNet. Companion Peace Corps Tunisian Arabic course. [open] (Internet Archive).
- Community Tunisian↔English dictionary, ~17,000 entries with example sentences and audio. Scraper: ArmelVidali/derja_ninja_scraper. [open] (web).
- Online Tunisian Arabic dictionary. [open] (web).
- Collaborative lexicon plus a Swadesh list. [open].
- Boujelbane et al., Building bilingual lexicon to create Dialect Tunisian corpora (WANLP 2013). [paper only].
- Sadat et al., TDA–MSA lexicon (COLING LG-LP 2014, ResearchGate). [paper only].
- MADAR Lexicon — multi-dialect, includes Tunis and Sfax entries. [on request].
- The dedicated Tunisian LLM instruction benchmark. Ben Hassine, Arrak, Addhoum, Wilson (CoCoA Lab, Univ. of Michigan-Flint), EMNLP 2025.
- 744 Tunisian instructions with gold human responses; ten Arabic-claiming LLMs scored by humans and an LLM judge on quality, correctness, relevance, dialectal adherence. Finding: most LLMs struggle in Tunisian.
- Repo: Souha-BH/TounsiBench-… (
data/= instructions, native gold responses, topic labels, evaluated model outputs; plus a GPT-4o evaluation notebook to score your own model). Paper: EMNLP 2025. [open].
- Parallel Tunizi ↔ standard Tunisian ↔ English corpus with sentiment labels; benchmarks LLMs on transliteration, translation, sentiment. Mahdi et al., 2025. [open] (arXiv).
- MMLU-style multiple-choice evaluation benchmark in Tunisian Derja, ~21.7k rows. LINAGORA, 2025. Built to evaluate the Labess LLM in-dialect.
- [open] (HuggingFace).
- Benchmark dataset for Tunisian/local dialects, LLM-eval oriented. Size/tasks not fully verified. [open] (HuggingFace).
The instruction / SFT / DPO / synthetic data layer behind the Tunisian LLMs in MODELS.md. Mostly 2025–2026, mostly open, and mostly unreviewed community data, quality varies, so treat sizes as raw counts.
- linagora/Tunisian_Derja_Dataset — 2.23M-row Derja corpus used for continual pre-training (also listed under raw corpora). [open].
- wghezaiel/SFT-Tunisian-Derja (~40k), wghezaiel/DPO-Tunisian-Derja (~44k), wghezaiel/SFT_derja_dataset (~48k), wghezaiel/derja_to_msa_dataset (~18k), wghezaiel/fw_tunisian_derja (~37k), wghezaiel/TunisianWikipedia-QA (~34k). Wajdi Ghezaiel (LINAGORA), 2025. The SFT/DPO alignment stack behind Labess. [open].
- ESPRIT-Derja-Instruct — ~7,013 bilingual Arabic/Arabizi Tunisian instruction pairs distilled from GPT-4o; trains the ESPRIT-Derja-Qwen3-8B model. [open] (bundled with the model repo).
- khaled123/Tunisian_Dialectic_English_Derja (~1.66M rows, Derja↔English, translations + sentiment + generation), khaled123/tunisiansynthinstract (~266k synthetic instructions), khaled123/Tunisianderjasynthtranslation (~436k), khaled123/tuniset (~29k). 2024. [open]. (Note: this account hosts many datasets of uneven quality — some are noisy dumps; verify before use.)
- Datasmartly/darja-tunisie-chat — Tunisian chat dataset, 2025. [gated] (must accept conditions; no card).
- abdouuu/tunisian_chatbot_data — ~1,426 instruction pairs, 2024. Weak as a Derja-generation set (questions in Derja, answers largely MSA/English). [open].
- Language resources for Maghrebi Arabic dialects' NLP: a survey (LREV 2020) — Younes, Souissi, Achour, Ferchichi. Catalogs 158 Maghrebi (incl. Tunisian) resources. The most thorough survey for Tunisian coverage.
- Maghrebi Arabic dialect processing: an overview (ICNLSSP 2017) — Harrat, Meftouh, Smaïli.
- Survey on Corpora Availability for the Tunisian Dialect (JCCO 2018) — Younes et al. Tunisian-specific.
- Critical description of Tunisian Arabic linguistic resources (ACLing 2018) — Mekki, Zribi, Ellouze, Belguith.
- Arabic natural language processing: an overview (2019) — Guellil et al.
- Natural Language Processing for Dialectal Arabic: A Survey (WANLP 2015) — Shoufan and Al-Ameri.
- A Survey on Dialect Arabic Processing and Analysis (ACM TALLIP 2025).
- Revisiting Common Assumptions about Arabic Dialects in NLP (ACL 2025) — Keleg, Goldwater, Magdy.
- Critical Survey of the Freely Available Arabic Corpora (2017) — Wajdi Zaghouani. Inventory of open Arabic corpora, dialectal ones included.
- NADI — Nuanced Arabic Dialect Identification (portal, GitHub). Tunisia is a target country each edition: 2020, 2021, 2022, 2023, 2024 (includes a Tunisian→MSA MT subtask). NADI 2025 and NADI 2026 added speech tracks — see SPEECH.md.
- MADAR Shared Task 2019 — city-level dialect ID (25 cities incl. Tunis and Sfax) + Twitter user dialect ID.
- IWSLT 2022 / 2023 dialectal speech translation — Tunisian↔English (see SPEECH.md).
- Arabic Dialect Identification (LREC 2018) — earlier DID benchmark with Tunisian-relevant data.
- mena-open-data/tunisian-dataset — HuggingFace collection of ~55 Tunisian datasets; the single richest live index found.
- linagora/tunisian-arabic-dialect-speech-and-text-modeling — LinTO/Labess ecosystem (datasets + models) in one collection.
- Definitive Guide of Tunisian Dialect NLP Resources — Chiraz Ben Abdelkader.
- Awesome-Tunisian-DATAI — TounesAI.
- ANLP-RG corpora page (MIRACL, Sfax).
- tunis-ai HF collections (datasets, models).
If this inventory is useful in your research, please cite it:
@misc{jlali2026tunisiannlp,
author = {Jlali, Fatma},
title = {Tunisian Arabic {NLP} Resources: A Curated Inventory},
year = {2026},
howpublished = {\url{https://github.com/jjlalli/Tunisian-NLP-Resources}},
note = {Living inventory of datasets, models, and papers for Tunisian Arabic (aeb)}
}A paper describing this inventory is in preparation; this entry will be updated when it is available.
The list itself is released under CC BY 4.0. The linked resources keep their own licenses — check each entry.
Maintained by Fatma Jlali. Contributions and corrections welcome — see CONTRIBUTING.md. Last compiled: July 2026.