Non-profit · Registered in The Gambia hello@anndal.org Support our work
Anndal Institute African Language Technology
Partner with us
Open language technology · Africa

Teaching machines to hear the languages they never learned.

Africa carries roughly a third of the world's languages and almost none of its language data. Anndal Institute builds the missing layer — open corpora, open models, and documentation good enough for anyone else to build on. We start with Mandinka, and we are not stopping there.

Open corpora/ Open models/ Speech-first languages/ Trilingual research/ CC BY-SA 4.0
~2,000
languages spoken in Africa — about a third of the world's total
37m+
speakers of the first language we are building for
2
research programmes running: language technology, and energy policy
100%
of what we produce released openly under CC BY-SA 4.0
Who we are

Africa speaks two thousand languages. Artificial intelligence understands a handful.

Anndal Institute is a non-profit research institute working on African language technology. We build the resources that have to exist before anything else can — speech corpora, text corpora, benchmarks and models — and we release all of them openly, so the next team does not start from zero.

The gap is not evenly distributed. A handful of African languages have usable data; most have almost none, and the ones that live mostly in speech rather than text are furthest behind. Those are the ones we go after first. Our current flagship builds for Mandinka; the method is designed to carry across the Manding continuum and beyond.

The same discipline — build the corpus properly, document it, release it — drives our second programme, on energy policy between Africa and Europe.

See our projects →

images/community.jpg
What we do

One method, applied wherever African public life is recorded.

We build corpora from the places a society talks about itself — parliaments, markets, ministries — and turn them into resources anyone can use. The engineering is the same whether the recording is a farmer describing a sick crop or a regional energy directive nobody has read end to end.

01 · Corpora

Speech and language corpora

Dual-register collection — the formal speech of parliament alongside the everyday speech of markets and farms, recorded in real background noise, with consent given in the speaker's own language.

02 · Models

Open models and tools

We fine-tune strong multilingual foundations rather than starting from zero, and report honest benchmarks against an unadapted baseline so others know exactly where a model works and where it fails.

03 · Policy

Policy and energy futures

The same corpus discipline turned on energy policy — every document hashed, archived and citable — analysed for where regional commitments and national plans quietly diverge.

04 · People

Skills that stay

Transcribers, field coordinators and student researchers — mostly young Africans — trained in language-data work. When a grant ends, the resource and the people who can extend it are both still here.

Projects

Two programmes running, and a method built to carry to the next language.

Each project is a corpus first and a model second. That order is deliberate: the corpus outlives the model, and it is what lets somebody else pick the work up.

images/assembly.jpg
Active Language technology

MandiAI

Open speech and language infrastructure for Mandinka and the Manding continuum — the first open speech recognition model for a language of 37 million people, trained on parliamentary and community speech at once.

ISO 639-3 · mnk50–100 hours12 months

Read the project →

images/energy.jpg
Active Policy research

Africa–Europe Energy Frontier

Trilingual corpora of African and European energy policy, built to a citable provenance standard and analysed for where regional commitments and national plans diverge — from ECOWAS renewable-energy targets to EU rules on imported hydrogen.

141 documents10,063 pagesEN · FR · PT

Read the project →

What comes next

The continuum, then the languages beside it

Mandinka sits at the western end of the Manding continuum, alongside Maninka in Guinea and Mali, Bambara in Mali, and Dyula in Côte d'Ivoire and Burkina Faso. They share vocabulary, grammar and sound systems, so a well-documented Mandinka corpus and pipeline is not an endpoint — it is a foundation that adapts to neighbouring varieties at a fraction of the cost of starting again. That expansion, and the other speech-first languages of the region, is what we are seeking partners for now.

Project 01 · Language technology

MandiAI

Open speech and language infrastructure for Mandinka and the Manding continuum — the first open speech recognition model for a language of 37 million people.

A model trained on Parliament and on the millet field

A model that only knows the careful speech of a debating chamber will stumble in a noisy market. So MandiAI draws on two corpora at once: Gambia National Assembly recordings, accessed under a memorandum of understanding with the Ministry of Information, and everyday Mandinka recorded across all seven regions with Clean Earth Gambia.

One stream teaches the model the language of the record. The other teaches it the language people actually use.

ISO 639-3 · mnk Speech recognition Dual register CC BY-SA 4.0
images/mandiai.jpg
Corpus target
50–100 transcribed hours across both registers, with region, setting and speaker metadata per hour.
Base models
Meta MMS and OpenAI Whisper, fine-tuned and compared. We publish the numbers behind whichever we release.
Evaluation
Word and Character Error Rate on a held-out set, per register, against an unadapted baseline — plus native-speaker review.
Release
Audio, transcripts, weights, code and model cards on HuggingFace, GitHub and the Masakhane repository.
Q1 · Months 1–3

Foundations

Data terms with the Ministry of Information, consent and compensation procedures, orthographic conventions fixed and documented.

Q2 · Months 4–6

Collection

Community recording across the seven regions, weighted toward agricultural speech. Transcription and quality review run in parallel.

Q3 · Months 7–9

Build

Corpus assembled. Fine-tuning experiments across base models and data mixes, with the test set held out from the start.

Q4 · Months 10–12

Release and return

Layered evaluation, open publication of everything, and results presented back to participating communities in Mandinka.

Project 02 · Policy research

Where Africa's energy transition meets Europe's rulebook.

Europe writes import rules for green hydrogen; West Africa writes regional renewable-energy targets. Both arrive as thousands of pages in three languages that almost nobody reads end to end. We turn them into structured, citable corpora and analyse where the commitments actually diverge.

Corpora in construction
Corpus Scope Documents Pages Languages
ECOWAS policy coherenceRegional renewable-energy and efficiency policy against the national plans of 15 member states, 2013–20251258,939EN · FR · PT
Africa–EU hydrogen frontierEU import rules for renewable hydrogen against the export strategies of five African states161,124EN · FR
Parliamentary discourseLongitudinal corpus of the German Bundestag, 1949–2026, and an in-progress corpus of the Gambia National Assembly1,033,723DE · EN · mnk

Every document carries its source URL, byte size, SHA-256 hash, page count and a web-archive snapshot. Several issuing ministries have already taken their own documents offline; the archived copy is what a reader will still be able to open in five years, and that is what we cite.

How we work

A language technology is only as good as its relationship with the people who speak the language.

Consent in the language being recorded

Explained and given in the speaker's own language, not buried in a form nobody reads. Sensitive material is handled under procedures agreed with local input.

Contributors are paid

Speech has real value. Treating it as free repeats exactly the extractive pattern this work exists to break, so every contributor is fairly compensated for their time.

Open by default, openly licensed

CC BY-SA 4.0 keeps the resource free and obliges anyone building on it to keep their additions open too. The resource belongs to its speakers, not to us.

Balanced, not accidentally narrow

A language is spoken differently by a young trader in the city and an elder farmer upriver. We track region, gender and age group as we collect, and balance deliberately.

Reproducible and provenanced

Pipelines are scripted rather than manual, so nothing depends on one person's memory and a neighbouring language variety can plug straight in later.

Results returned to the community

We close the loop where we opened it — going back to the communities that gave us their voices and presenting what those voices built, in their own language.

Our team

Three leads, all Gambian, and the team the work builds around them.

AN

Adama Njie

Director · Corpora & NLP

PhD researcher at RWTH Aachen University working on data-driven policy systems and natural language processing. Built a longitudinal corpus of 1,033,723 German Bundestag speeches and is constructing the Gambia National Assembly parliamentary corpus. MSc Advanced Computer Science, Cardiff University, Chevening Scholar.

P2

Partner 2

Community & field data

Leads community engagement, field recording across all seven regions, informed consent given in the language being recorded, and contributor relations. Named once confirmed.

P3

Partner 3

Data architecture & evaluation

Leads corpus processing, storage architecture, quality control and the evaluation framework, including benchmarking and error analysis. Named once confirmed.

Partners & collaborators

Rooted in The Gambia, across civil society, academia and government.

Community

Clean Earth Gambia

UNCCD-accredited environmental organisation present in all seven regions. Leads community engagement, field recording, informed consent and contributor relations.

Institutional data

Ministry of Information & the National Assembly

Provide access to National Assembly recordings and transcripts under a memorandum of understanding — the formal register of the corpus.

Academic

University of The Gambia, School of ITC

Research grounding, local technical capacity, and a route for student involvement and knowledge transfer.

Network

Masakhane & the African NLP community

We publish into the community's shared repositories and build on its participatory research tradition rather than working around it.

Partner with us

Three ways to work with us, all of which we can start within a quarter.

We work with funders, research groups and institutions in Africa, Europe and North America. Tell us which of these is closest and we will send the detail.

Research funders

Fund a language from zero to first model

A full twelve-month programme — corpus, model, open release and community return — for one language costs in the order of US$100,000, plus compute.

  • Data collection and fair contributor payment
  • Transcription and quality assurance
  • Training, benchmarking and open release
Request the budget
Technical partners

Contribute compute or engineering

Fine-tuning multilingual speech foundations is the cost we cannot absorb locally. Cloud credits, GPU access or ML engineering time all move this directly.

  • GPU hours for fine-tuning experiments
  • Guidance on low-resource speech evaluation
  • Corpus and checkpoint hosting
Talk to us
Institutions

Partner on a language or a corpus

Universities, parliaments, broadcasters and NGOs across West Africa hold the recordings that make the next model possible.

  • Joint research and co-authorship
  • Archive and broadcast data agreements
  • Student placements and training
Propose a partnership
Support our work

Every hour of transcribed speech costs about US$135.

Transcription, a second-pass quality review, metadata annotation and fair payment to the speaker. Fifty of those hours is a working model. It is the most direct thing anyone outside the region can fund.

InstitutionsGrant agreements and restricted project funding, invoiced from the registered entity.
IndividualsBank transfer today; card giving is being set up. Write to us and we will send details.
In kindCompute credits, recording equipment, transcription software licences.
Contact

Tell us what you are building, and we will tell you honestly whether we can help.

We answer every serious enquiry from a funder, a research group or an institution holding language data. If you speak one of the languages we work on and want to contribute recordings, write to us too — that is the part of this work that cannot be bought.

Before this page goes live: register the domain, replace the placeholder GitHub and HuggingFace links above, set the form endpoint in js/config.js, and name the two team members once they have agreed to appear. Partner organisations are confirmed. Delete this note when the list is empty.

We reply to every enquiry within a week. Nothing you send here is shared outside the three people named on this page.