Datasets

Language

If automated speech recognition (ASR) systems are only built for a small percentage of international languages, the result is a pronounced gap in access to the digital world that is glaring in low- and middle-income contexts globally. Translation, speech recognition, and more datasets enable the promise of digital communication for all.

Language dataset imagery
Released datasets
24
Contributing authors
32+
Access & licensing
Open
Domain area
Language
Why it matters

Language and machine learning

Lacuna Fund language datasets create openly accessible text and speech resources that fuel natural language processing technologies in diverse languages across low- and middle-income contexts globally. Explore and download released datasets below.

Locally owned by design

Every dataset here was produced by teams in low- and middle-income contexts and is openly accessible to the international community.
Spotlight

Featured dataset

Recently listed

A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

This dataset is the first large-scale human-annotated Twitter sentiment dataset for Hausa, Igbo, Nigerian-Pidgin, and Yorùbá, the four most widely spoken languages in Nigeria.
Released datasets

24 datasets

Search the collection, then expand an entry to see authors, descriptions and access links.

Showing 24 of 24