Speech and language corpora
Dual-register collection — the formal speech of parliament alongside the everyday speech of markets and farms, recorded in real background noise, with consent given in the speaker's own language.
Africa carries roughly a third of the world's languages and almost none of its language data. Anndal Institute builds the missing layer — open corpora, open models, and documentation good enough for anyone else to build on. We start with Mandinka, and we are not stopping there.
Anndal Institute is a non-profit research institute working on African language technology. We build the resources that have to exist before anything else can — speech corpora, text corpora, benchmarks and models — and we release all of them openly, so the next team does not start from zero.
The gap is not evenly distributed. A handful of African languages have usable data; most have almost none, and the ones that live mostly in speech rather than text are furthest behind. Those are the ones we go after first. Our current flagship builds for Mandinka; the method is designed to carry across the Manding continuum and beyond.
The same discipline — build the corpus properly, document it, release it — drives our second programme, on energy policy between Africa and Europe.
We build corpora from the places a society talks about itself — parliaments, markets, ministries — and turn them into resources anyone can use. The engineering is the same whether the recording is a farmer describing a sick crop or a regional energy directive nobody has read end to end.
Dual-register collection — the formal speech of parliament alongside the everyday speech of markets and farms, recorded in real background noise, with consent given in the speaker's own language.
We fine-tune strong multilingual foundations rather than starting from zero, and report honest benchmarks against an unadapted baseline so others know exactly where a model works and where it fails.
The same corpus discipline turned on energy policy — every document hashed, archived and citable — analysed for where regional commitments and national plans quietly diverge.
Transcribers, field coordinators and student researchers — mostly young Africans — trained in language-data work. When a grant ends, the resource and the people who can extend it are both still here.
Each project is a corpus first and a model second. That order is deliberate: the corpus outlives the model, and it is what lets somebody else pick the work up.
Open speech and language infrastructure for Mandinka and the Manding continuum — the first open speech recognition model for a language of 37 million people, trained on parliamentary and community speech at once.
Trilingual corpora of African and European energy policy, built to a citable provenance standard and analysed for where regional commitments and national plans diverge — from ECOWAS renewable-energy targets to EU rules on imported hydrogen.
Mandinka sits at the western end of the Manding continuum, alongside Maninka in Guinea and Mali, Bambara in Mali, and Dyula in Côte d'Ivoire and Burkina Faso. They share vocabulary, grammar and sound systems, so a well-documented Mandinka corpus and pipeline is not an endpoint — it is a foundation that adapts to neighbouring varieties at a fraction of the cost of starting again. That expansion, and the other speech-first languages of the region, is what we are seeking partners for now.
Open speech and language infrastructure for Mandinka and the Manding continuum — the first open speech recognition model for a language of 37 million people.
A model that only knows the careful speech of a debating chamber will stumble in a noisy market. So MandiAI draws on two corpora at once: Gambia National Assembly recordings, accessed under a memorandum of understanding with the Ministry of Information, and everyday Mandinka recorded across all seven regions with Clean Earth Gambia.
One stream teaches the model the language of the record. The other teaches it the language people actually use.
Data terms with the Ministry of Information, consent and compensation procedures, orthographic conventions fixed and documented.
Community recording across the seven regions, weighted toward agricultural speech. Transcription and quality review run in parallel.
Corpus assembled. Fine-tuning experiments across base models and data mixes, with the test set held out from the start.
Layered evaluation, open publication of everything, and results presented back to participating communities in Mandinka.
Europe writes import rules for green hydrogen; West Africa writes regional renewable-energy targets. Both arrive as thousands of pages in three languages that almost nobody reads end to end. We turn them into structured, citable corpora and analyse where the commitments actually diverge.
| Corpus | Scope | Documents | Pages | Languages |
|---|---|---|---|---|
| ECOWAS policy coherence | Regional renewable-energy and efficiency policy against the national plans of 15 member states, 2013–2025 | 125 | 8,939 | EN · FR · PT |
| Africa–EU hydrogen frontier | EU import rules for renewable hydrogen against the export strategies of five African states | 16 | 1,124 | EN · FR |
| Parliamentary discourse | Longitudinal corpus of the German Bundestag, 1949–2026, and an in-progress corpus of the Gambia National Assembly | 1,033,723 | — | DE · EN · mnk |
Every document carries its source URL, byte size, SHA-256 hash, page count and a web-archive snapshot. Several issuing ministries have already taken their own documents offline; the archived copy is what a reader will still be able to open in five years, and that is what we cite.
Explained and given in the speaker's own language, not buried in a form nobody reads. Sensitive material is handled under procedures agreed with local input.
Speech has real value. Treating it as free repeats exactly the extractive pattern this work exists to break, so every contributor is fairly compensated for their time.
CC BY-SA 4.0 keeps the resource free and obliges anyone building on it to keep their additions open too. The resource belongs to its speakers, not to us.
A language is spoken differently by a young trader in the city and an elder farmer upriver. We track region, gender and age group as we collect, and balance deliberately.
Pipelines are scripted rather than manual, so nothing depends on one person's memory and a neighbouring language variety can plug straight in later.
We close the loop where we opened it — going back to the communities that gave us their voices and presenting what those voices built, in their own language.
PhD researcher at RWTH Aachen University working on data-driven policy systems and natural language processing. Built a longitudinal corpus of 1,033,723 German Bundestag speeches and is constructing the Gambia National Assembly parliamentary corpus. MSc Advanced Computer Science, Cardiff University, Chevening Scholar.
Leads community engagement, field recording across all seven regions, informed consent given in the language being recorded, and contributor relations. Named once confirmed.
Leads corpus processing, storage architecture, quality control and the evaluation framework, including benchmarking and error analysis. Named once confirmed.
UNCCD-accredited environmental organisation present in all seven regions. Leads community engagement, field recording, informed consent and contributor relations.
Provide access to National Assembly recordings and transcripts under a memorandum of understanding — the formal register of the corpus.
Research grounding, local technical capacity, and a route for student involvement and knowledge transfer.
We publish into the community's shared repositories and build on its participatory research tradition rather than working around it.
We work with funders, research groups and institutions in Africa, Europe and North America. Tell us which of these is closest and we will send the detail.
A full twelve-month programme — corpus, model, open release and community return — for one language costs in the order of US$100,000, plus compute.
Fine-tuning multilingual speech foundations is the cost we cannot absorb locally. Cloud credits, GPU access or ML engineering time all move this directly.
Universities, parliaments, broadcasters and NGOs across West Africa hold the recordings that make the next model possible.
Transcription, a second-pass quality review, metadata annotation and fair payment to the speaker. Fifty of those hours is a working model. It is the most direct thing anyone outside the region can fund.
We answer every serious enquiry from a funder, a research group or an institution holding language data. If you speak one of the languages we work on and want to contribute recordings, write to us too — that is the part of this work that cannot be bought.