A curated list of resources for the conservation, development, and documentation of low resource (human) languages.
-
Updated
Jun 26, 2026 - TeX
A curated list of resources for the conservation, development, and documentation of low resource (human) languages.
This repository contains the code, data, and models of the paper titled "XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages" published in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021.
[EMNLP 2023] 💬 Language Identification with Support for More Than 2000 Labels
This repository contains the code and data of the paper titled "Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for Bengali-English Machine Translation" published in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), November 16 - November 20, 2020.
A repository for publicly/freely available Natural Language Processing (NLP) datasets for African languages.
Syllable-aware BPE tokenizer for the Amharic language (አማርኛ) – fast, accurate, trainable.
The open, machine-readable standard for the Somali language — orthography, grammar, terminology, translation, and AI/benchmark resources, as a versioned, citable RFC-style catalog.
NLP pipelines for Tagalog using spaCy
[PRL 2025, APSIPA 2022] Syllable Analysis Data Augmentation (SADA), This project introduces a glyph dictionary and grammar-aware augmentation strategy designed to enhance Khmer palm leaf manuscript recognition. By modeling the language's grammatical structure, we support more robust OCR performance in low-resource settings.
Speech synthesis (TTS) in low-resource languages by training from scratch with Fastpitch and fine-tuning with HifiGan
Open-source benchmark datasets and pretrained transformer models in the Filipino language.
SomNLP-Corpus is an open, high-quality Somali text corpus designed to support research and development in Natural Language Processing (NLP), Large Language Models (LLMs), and other AI applications for the Somali language.
SemEval2024-task 11: Bridging the Gap in Text-Based Emotion Detection
CogNet: a large-scale, high-quality cognate database for 338 languages, 1.07M words, and 8.1 million cognates
The EveryVoice TTS Toolkit - Text To Speech for your language
Tiny language detector + tokenizer for 50+ SE/South Asian languages (Burmese, Karen, Chin, Shan, Mon, Khmer, Lao, Thai, Tamil, Hindi, …). Revamp of pyidaungsu.
This is an ASR corpus for Bemba language. It contains read speech from diverse publicly available Bemba sources; Literature Books, Radio/TV shows transcripts, Youtube Video transcripts, Online sources. The corpus has 14, 438 utterances culminating into over 24 hours of speech.
This is a repository for NaijaSenti. A Lacuna Funded Project for the development of sentiment corpus for four Nigerian languages: Igbo, Hausa, Yoruba and Pidgin.
[ACL'24] MC^2: A Multilingual Corpus of Minority Languages in China (Tibetan, Uyghur, Kazakh, and Mongolian)
Curated list of publicly available parallel corpus for Indian Languages
Add a description, image, and links to the low-resource-languages topic page so that developers can more easily learn about it.
To associate your repository with the low-resource-languages topic, visit your repo's landing page and select "manage topics."