A community project in Ebonyi State, Nigeria

Voices that have
never been heard.

We are building the first open speech and text dataset for the Izzi language cluster. Izzi, Ezaa, Ikwo, and Mgbo together have millions of speakers in Ebonyi State. None of them have a single hour of transcribed audio in any public archive.

The gap in numbers

Four figures that explain the project

517

languages in Nigeria

of them alive, spoken, needed

Ethnologue count for Nigeria
4

with meaningful NLP work

under one percent of the total

NaijaNLP survey, arXiv 2025
0

datasets for the Izzi cluster

no corpus, no ASR, no benchmark

Verified absence, Izi language - Wikipedia
2.4%

of new African voice data from speakers 50+

elder voices are nearly absent

Lanfrica analysis of African Next Voices

What Olubridge is

We are recording speech and collecting text in Izzi, Ezaa, Ikwo, and Mgbo, four languages of Ebonyi State that share a cluster but sit outside Central Igbo. No dataset exists for any of them. We are building the first.

Read the background

Why now

Only 2.4% of new African voice data comes from speakers older than fifty. The generation that holds the deepest vocabulary is still here. Voice AI trained without them will fail the people who need it most.

See the landscape

How to help

You can record a session, fund a recording day, or connect us with a community leader. Each hour of clean elder speech is a data point that does not exist anywhere else.

Get involved
“Onye wetara oji wetara ndu.”

The one who brings kola nut brings life.

Igbo proverb, shared across the Izzi cluster