Lacuna Fund language datasets create openly accessible text and speech resources that fuel natural language processing technologies in diverse languages across low- and middle-income contexts globally. Explore and download released datasets below.
- Released datasets
- 24
- Contributing authors
- 32+
- Access & licensing
- Open
- Domain area
- Language
Why it matters
Language and machine learning
Locally owned by design
Every dataset here was produced by teams in low- and middle-income contexts and is openly accessible to the international community.
Spotlight
Featured dataset
Recently listed
A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis
This dataset is the first large-scale human-annotated Twitter sentiment dataset for Hausa, Igbo, Nigerian-Pidgin, and Yorùbá, the four most widely spoken languages in Nigeria.
Released datasets
24 datasets
Search the collection, then expand an entry to see authors, descriptions and access links.
Showing 24 of 24
