About this report
Atilingo is a nonprofit technology organization. Its specific purpose is to operate exclusively for educational, scientific and charitable purposes: preserving, documenting and revitalizing Indigenous, endangered and underrepresented languages and the cultural knowledge and heritage associated with them; developing and supporting educational programs, digital tools and research that advance language learning, linguistic diversity and cultural preservation; and creating, developing and disseminating open, accessible resources and technologies that support the study, use and preservation of underrepresented languages. We work with speaker communities to record oral material, transcribe it with AI assistance and native-speaker review, and publish it as structured data that researchers, developers and the communities themselves can use.
This is the first in a planned annual series, setting out what is currently known — and, more importantly, what is not known — about the number, distribution, vitality and digital representation of African languages, as a baseline funders, ministries, researchers and technologists can cite from a common footing.
Evidence and method
Findings here draw on three kinds of evidence: established results from published research and language catalogs, estimates from incomplete or inconsistent data, and Atilingo’s own interpretations of patterns across sources. Where the evidence is uneven, we label the confidence — ranges, qualifications and “unknown” classifications are findings, not gaps.
The four principal reference works are Glottolog, Ethnologue, the UNESCO World Atlas of Languages, and the Catalog of Endangered Languages (ELCat), supplemented by national statistical returns and the published computational-linguistics literature. See Section X for method and the Appendix for the full enumeration.
Executive summary
Seven findings, the thesis, and the recommendations at a glance.
Africa holds roughly a third of the world’s languages and a small fraction of the world’s language data. The conventional framing of language loss as a preservation problem — communities losing speakers — is accurate but incomplete.
The more urgent emergency: the overwhelming majority of African languages are absent from the digital and computational record, and therefore absent from the technologies that will mediate education, government, commerce, and knowledge for the next century.
Seven findings
No agreed count exists, and the disagreement is the finding
Published estimates for Africa range from about 1,250 to over 3,000. The catalogs diverge most precisely where survey coverage is thinnest, making the spread a usable proxy for how poorly documented a region is.
The most-cited endangerment data is sixteen years old
UNESCO’s Atlas dataset was completed in 2010 and has not been substantively updated. Every widely quoted African endangerment figure predates the mobile-internet and AI era it is now used to describe.
Several hundred languages are at risk, and up to a tenth may not survive the century
UNESCO records nearly 428 African languages as threatened; an independent survey identified over 500 as highly endangered, extinct or close to it. Both assessments are now more than fifteen years old.
Digital invisibility affects almost every African language, not only endangered ones
The majority of African languages have no annotated digital data of any kind. Fewer than two per cent meet the report’s higher digital-resource threshold — defined in the methodology as broad coverage across multiple computational tasks and usable text and/or speech resources — comparable to that available for a typical mid-sized European language.
The exclusion begins in the research pipeline
In 2018, five of 2,695 author affiliations at the five major NLP conferences were African institutions. Resources for African languages will not appear as a by-product of the field’s normal activity.
Linguistic density and documentation are inversely related
West Africa holds roughly 890 living languages, more than any comparable region on earth, and is among the least surveyed. Central Africa shows the same pattern more acutely.
Sign languages, creoles and urban vernaculars are systematically uncounted
African sign languages, contact languages such as Nigerian Pidgin with tens of millions of speakers, and fast-growing urban vernaculars sit almost entirely outside the reference catalogs.
Recommendations at a glance
- Fund a continental vitality reassessment; the current evidence base is sixteen years old.
- Add or restore language questions to national census instruments, with published methodology.
- Treat orthographic standardization and keyboard/input support as digital infrastructure rather than cultural policy.
- Recognize mother-tongue instruction as a data problem as much as a pedagogical one — materials require corpora.
- Fund audio-first, consented, openly licensed corpora rather than text-only or extractive collection.
- Fund African research institutions and participatory community programs directly, not commissioned deliverables.
- Set breadth targets, not depth targets: coverage across many languages beats depth in a few.
- Require open licensing and community data sovereignty as a condition of funding, and fund maintenance, not only creation.
Introduction
Why this report exists, what it argues, and what it deliberately does not claim.
1. Why this report, and why now
Two developments make 2026 a reasonable moment to restate the position of Africa’s languages. The first: the most widely cited endangerment data for the continent has aged badly. UNESCO’s Atlas of the World’s Languages in Danger dataset was completed in 2010 and has not been substantively updated since, so the figures press coverage, funding proposals and policy documents reproduce describe a situation more than fifteen years old.
The second: a new and largely unmeasured form of language loss has emerged in that same interval. Between 2010 and 2026, the primary interfaces for information, government services, commerce, and education shifted decisively toward systems built on large text and speech corpora. Those corpora do not merely underserve the languages absent from them; they cannot see those languages at all. A language with two million healthy speakers and no digital footprint is, from the perspective of a search engine, a translation system or a speech interface, indistinguishable from a language with no speakers at all.
How many African languages there are, how they are faring, and to what extent those languages exist in the digital record are inseparable questions. The first is the traditional subject of language documentation. The second determines whether the answer to the first will still be true in fifty years.
2. The argument
Our thesis posits that the emergency facing African languages has taken on a second, increasingly significant dimension: digital invisibility. While the loss of speakers remains the most immediate threat to the survival of critically endangered languages, the lack of digital presence now affects a much broader range of languages, including those with substantial and active speaker communities. Consequently, the central question has evolved; it is no longer solely about whether a language has speakers, but rather whether it is represented in the digital infrastructures that mediate knowledge, education, public services, and communication. This is not an argument that endangerment does not matter. Several hundred African languages have small, aging speaker populations and will not survive the century without intervention, and the loss of each is irreversible. It argues instead about where the leverage now lies and how large the affected population is.
Classical endangerment affects a minority of African languages — those with the smallest and most isolated speaker communities. Digital invisibility affects almost all of them, including the largest and healthiest. Hausa, Yoruba, Amharic, Oromo, Igbo and Zulu together account for hundreds of millions of speakers, and all of them remain poorly represented in the systems that now mediate access to information. For the several hundred languages with fewer than ten thousand speakers, digital invisibility and endangerment are the same problem, because a language with no recorded material has no route to revitalization once transmission stops.
In practice, priorities must reorder. Documentation has historically prioritized urgency: record the most endangered languages first. That logic remains sound for the critical cases, but it produces an archive optimized for scholarship rather than for use. An approach organized around usability — machine-readable, openly licensed, audio-anchored data across the full range of languages — serves both the critical cases and the large, healthy languages simultaneously, and it is the only approach that scales to 2,000 languages.
3. Limits
No new primary fieldwork underlies this synthesis of existing catalogs, statistical returns, and published literature, and its accuracy is bounded accordingly. No single authoritative count of African languages exists, so none is offered here; the internal linguistic structure of the languages discussed goes unassessed. And technology does not substitute for intergenerational transmission: a language survives because children speak it, not because it has a dataset.
We also note our own position: Atilingo builds language datasets, and a report arguing that language datasets are urgently needed is not disinterested. We manage this by grounding every claim in third-party sources, confining our own work to a single short section, and publishing the underlying data openly so the conclusions can be checked against it.
How many languages does Africa have?
Four catalogs, four answers, and what the disagreement between them measures.
No agreed answer exists. Published estimates for the number of living languages in Africa range from about 1,250 to over 3,000, and the range is not the result of careless scholarship. It reflects genuine, unresolved disagreement about what counts as a distinct language, compounded by uneven survey coverage across the continent.
1. What the four principal catalogs say
The figures below are the continental totals implied by the four reference works drawn on here. They are not directly comparable, and the reasons they are not comparable are more informative than the numbers themselves.
| Reference work | Continental total | Unit counted | Access |
|---|---|---|---|
| Glottolog | ~2,300–2,400 | Genealogically distinguished languoids, split where published literature supports it | Open, CC BY, fully downloadable |
| Ethnologue | ~2,140–2,150 | Living languages; functional criterion rooted in translation need; ISO 639-3 registrar | Subscription; licensing restricts republication |
| UNESCO WAL | 428 listed | Only languages assessed as endangered; silent on non-endangered languages | Open; core dataset completed 2010 |
| ELCat | Partial coverage | Endangered languages with structured vitality evidence and sources | Open, community-updatable |
Table 1 — Continental language totals by reference work. Totals are approximate and change between editions; see Section X for edition-specific figures and retrieval dates.
2. Why the numbers differ
The language/dialect boundary is partly political. Whether two neighboring varieties are counted as one language or two frequently depends on whether their speakers are treated as one people or two. States, churches, publishers and census authorities make that determination, not linguists. Splitting a dialect continuum into separate entries raises the count; merging it lowers the count. Neither operation is neutral, and both have been used to political ends — to inflate the apparent unity of a nation, or to fragment a minority.
The catalogs have different purposes.Ethnologue’s lineage is in Bible translation, and its unit of interest has historically been the community requiring a separate translation — a functional criterion. Glottolog is built for genealogical classification and bibliographic reference, and it splits where the published literature supports a split. UNESCO’s Atlas exists to flag endangerment, so it enumerates only languages assessed as at risk and is silent on the rest.
Coverage is geographically uneven. The Bantu-speaking areas of Central Africa, parts of South Sudan, Chad and the Ethiopian periphery contain languages that have never been the subject of a published grammar or wordlist. Their existence is attested, sometimes only by a single survey visit decades ago, but their boundaries and speaker numbers are effectively unknown.
Systematic exclusions. Three categories are routinely under-counted. African sign languages — of which several dozen are attested, some with substantial user communities — are absent from most continental figures. Widely spoken contact languages and creoles are inconsistently treated: Nigerian Pidgin, with tens of millions of speakers, is frequently omitted from national language counts altogether. And varieties of Arabic and Amazigh are grouped and split inconsistently between sources, moving the North African total by dozens.
3. The figure used here
Where a single number is required: approximately 2,000 to 2,200 living languages, the band supported by the two catalogs with full continental coverage and open or documented methodology. The true figure depends on decisions that are not purely linguistic; per-language data is in the Appendix for readers who want to apply their own splitting criteria.
By this estimate, Africa holds approximately one-third of the world’s living languages — the most linguistically diverse continental region by number of languages.
The linguistic landscape
Families, regions, writing systems, and the categories the count usually omits.
Africa’s languages are conventionally assigned to four large groupings, a scheme proposed in the 1960s that has not survived subsequent scrutiny intact. Two of the four are now seriously contested. The classification below reflects current consensus where it exists. It marks the disputes where it does not, because the dispute has practical consequences: a language whose family placement is unknown is typically also a language whose grammar is unpublished.
1. Families
Niger-Congo
~1,400–1,500 languages · c. 600 million speakersThe largest language family on earth by number of languages, spanning West, Central, East and Southern Africa. Its Bantu branch alone accounts for around 500 languages, including Swahili, Zulu, Shona and Kikuyu; its West African branches include Yoruba, Igbo, Akan and Fula.
Status of the classification: The family is widely accepted, but internal structure is unsettled, and the coherence of some proposed branches remains disputed. Many individual Bantu languages have no published grammar.
Afro-Asiatic
~375 languages · c. 500 million speakersSpans North Africa, the Sahel and the Horn, and extends into the Middle East. Includes Arabic varieties, Amharic, Tigrinya, Somali, Oromo, Hausa and the Amazigh (Berber) languages.
Status of the classification: Well established as a family. Disputes concern the treatment of Arabic and Amazigh varieties — whether to count them as single languages or as dozens — which moves the continental total substantially.
Nilo-Saharan
~200 languages · c. 60 million speakersA geographically dispersed grouping across the Sahel, Nile valley and Great Lakes, including Luo, Dinka, Nuer, Kanuri, Maasai and the Nubian languages.
Status of the classification: Contested. Many specialists regard Nilo-Saharan as a convenience grouping of several unrelated families rather than a demonstrated genetic unit. Nilo-Saharan is the least securely classified major grouping on the continent.
Khoe-Kwadi, Tuu, Kx’a
~30 languages · under 500,000 speakersSouthern African languages formerly bundled as “Khoisan”, including N|uu, Ju|’hoan, Khoekhoe and Naro. Distinguished by exceptionally large phonemic inventories, including extensive click systems.
Status of the classification:The single “Khoisan” family is now generally rejected; it is instead treated as three unrelated families plus isolates. They contain the most critically endangered languages in Africa, several of which have fewer than a dozen fluent speakers.
Austronesian and Indo-European
~10 languages · c. 40 million speakersMalagasy, an Austronesian language spoken across Madagascar by some 25 million people, and Afrikaans, an Indo-European language with roots in Dutch, alongside the colonial languages that remain official in most African states.
Status of the classification:Uncontroversial. Included here because these languages are frequently omitted from “African languages” counts despite being first languages for tens of millions of Africans.
Contact languages, creoles and sign languages
Several dozen attested · tens of millions of usersNigerian Pidgin, Krio, Cape Verdean Creole, Sango, Sheng and Camfranglais, alongside several dozen African sign languages from nationally recognized to village-scale.
Status of the classification: Inconsistently cataloged or absent entirely. It is the fastest-growing and least documented category of African language, and the one where the gap between speaker numbers and digital presence is widest.
2. Regional profiles
Linguistic density is not evenly spread. West Africa alone accounts for close to half of the continent’s languages, and Nigeria alone for around a quarter of West Africa’s total. The figures below are indicative midpoint estimates; per-country details are in the Appendix.
| Region | Languages | Principal families | Dominant pressures |
|---|---|---|---|
| West Africa | ~890 | Niger-Congo, Afro-Asiatic, Nilo-Saharan | Urban lingua francas; English/French in education |
| Central Africa | ~680 | Niger-Congo (Bantu), Nilo-Saharan | Weakest documentation; spread of French, Lingala and Sango |
| East Africa | ~440 | Niger-Congo (Bantu), Nilo-Saharan, Afro-Asiatic | Swahili as regional lingua franca; English in schooling |
| Horn of Africa | ~180 | Afro-Asiatic, Nilo-Saharan | Amharic and Somali dominance; conflict displacement |
| Southern Africa | ~130 | Niger-Congo (Bantu), Khoe-Kwadi, Tuu, Kx’a | Most critical endangerment cases; English/Afrikaans shift |
| North Africa | ~60 | Afro-Asiatic (Arabic, Amazigh) | Arabic diglossia; state policy toward Amazigh varieties |
Table 2 — Regional distribution. Counts are mid-point estimates across catalogs and should be read as orders of magnitude, not precise totals.
3. Writing systems
Most African languages that are written at all use a Latin-based orthography, frequently one devised by missionaries or colonial administrations and never formally standardized. A minority use Ge’ez script (Amharic, Tigrinya, Tigre), Arabic script (including the Ajami traditions used for Hausa, Wolof, Fula and Swahili), or one of the indigenous scripts developed in the nineteenth and twentieth centuries — Vai, N’Ko, Tifinagh, Osmanya, Adlam, Mwangwego and others.
Orthographic instability has a direct digital cost. Where a language has no agreed spelling standard, text corpora fragment across competing conventions and every downstream computational task degrades. Where a language uses tone marking that keyboards do not readily produce, writers omit the diacritics, and the resulting text is ambiguous in ways that are difficult to recover automatically. A substantial share of the digital deficit documented in Section VI is not a shortage of speakers willing to write but an absence of the standards and input methods that would make their writing usable.
Orthographic instability is also the strongest argument for treating audio as the primary medium. For a language with no orthographic standard, a recording is unambiguous data; a transcript is a set of decisions. Audio-first documentation defers the question of standardization rather than being blocked by it.
4. Categories the count usually omits
Sign languages. Several dozen African sign languages are attested, ranging from national languages with formal recognition to village sign languages with a few dozen users. They are absent from most continental counts, almost absent from digital resources, and, by construction, excluded from any documentation program built on audio.
Contact languages and urban vernaculars. Nigerian Pidgin, Sheng, Camfranglais, Tsotsitaal and similar varieties have very large and growing user populations, are the dominant medium of daily life for many urban Africans, and sit almost entirely outside the reference catalogs. They are simultaneously the fastest-growing and least documented category on the continent.
Languages of recent record only.A residual set of languages appears in the catalogs on the strength of a single twentieth-century survey and has not been verified since. For these, the honest status is not “endangered” or “vital” but unknown — and unknown status is itself a finding that current classification schemes have no way to express.
Vitality, endangerment and extinction
What the assessments record, what drives shift, and why the evidence base has aged badly.
Africa’s endangerment profile differs from that of the Americas, Australia or Siberia. Proportionally fewer African languages are moribund, because a large number of mid-sized languages remain in vigorous daily use across multiple generations. Some read that as reassurance. It should not be: the mechanisms that produced rapid language loss elsewhere are present and intensifying, and the data available to monitor them is worse than for any other continent.
1. What the assessments record
The principal published assessments are broadly consistent in magnitude while differing in method and vintage.
| Assessment | Vintage | What it records |
|---|---|---|
| UNESCO Atlas | Dataset completed 2010 | Records nearly 428 African languages as threatened across four tiers and warns that up to 10 per cent of African languages may vanish within a century. Not substantively updated since completion. |
| UNESCO Atlas (earlier edition) | c. 2002 | Of the roughly 1,400 African languages then cataloged, between 500 and 600 were assessed as in decline and 250 as under immediate threat. The same edition described Africa as linguistically the least-known continent. |
| Batibo continental survey | 2005 | A country-by-country survey identified 509 African languages as highly endangered, extinct or on the brink of extinction — broadly corroborating the UNESCO magnitude by an independent method. |
| Documented extinctions | Cumulative | At least 74 languages that diverged from their parent language in Africa are recorded as having no remaining native speakers and no spoken descendants. The actual figure is likely higher, because extinctions in poorly surveyed areas may go unrecorded. |
Table 3 — Principal published vitality assessments for Africa.
2. Drivers of shift
Education in ex-colonial languages. In most African states, the language of instruction shifts to English, French, Portuguese or Arabic within the first few years of primary school, marking the home language as inadequate for formal knowledge at the exact age when children are calibrating which language carries status.
Urbanization and regional lingua francas. Rural-to-urban migration places speakers of many languages in shared space, where a lingua franca — Swahili, Hausa, Wolof, Lingala, Amharic, Arabic — becomes the medium of daily life. In some communities, such processes can produce substantial intergenerational shift within two generations, although the pace varies considerably by language, region, domain of use, and community.
Absence of official status. A language with no role in government, courts, media or commerce offers its speakers no instrumental reason to maintain it. Official-language policy in most African states recognizes one or two languages out of dozens, and the unrecognized majority carry the full cost of that choice.
Small and dispersed speaker populations. Several hundred African languages have fewer than 10,000 speakers, and for many the community is not geographically concentrated. Below a certain population and density, ordinary demographic pressure is sufficient to end transmission without any specific hostile policy.
Digital absence. Treated at length in Section VI. Since the last continental vitality assessment, digital absence has moved from a marginal factor to a primary one, and it is the only driver on this list that current assessment instruments do not measure at all.
3. A note on measuring vitality
Speaker counts are the weakest link in the chain. Most published figures for African languages derive from national censuses, and censuses ask about language inconsistently, infrequently, or not at all. Some states do not collect language data because the results would be politically inconvenient. Others report figures unchanged across successive editions, which cannot be accurate for a continent whose population is growing rapidly. A significant number of the speaker estimates in circulation are extrapolations from decades-old surveys, carried forward and re-cited until their provenance is lost.
A speaker count alone also captures vitality poorly. A language with a million speakers, none of them under thirty, is in more danger than a language with five thousand speakers being raised as children’s first language. The standard assessment scales attempt to capture that dynamic through transmission-based criteria, but applying them requires community-level observation that has not been carried out for most African languages. For a substantial proportion of the continent’s languages, no current vitality assessment exists — and that omission is itself the honest assessment.
Digital invisibility
Absence from the data is absence from the century.
1. The shape of the deficit
African languages are near-absent from the large web-crawled corpora on which contemporary language technology is built. Their representation in Wikipedia, and in web crawls generally, is a small fraction of a per cent. Because those crawls are the raw material for translation systems, speech recognition and large language models, the absence propagates through every downstream system. Computational-linguistics literature describes the resulting condition as a digital language divide, in which standard processing pipelines simply fail to generalize to African linguistic contexts.
The widely used resource taxonomy in that literature places languages on a six-level scale based on the availability of labeled and unlabelled data. The great majority of African languages sit at the bottom two levels — the categories informally named the left-behinds and the scraping-bys — with no annotated data of any kind. A handful, notably Swahili, Hausa, Amharic and Afrikaans, reach the middle of the scale.
Digital-resource status is multidimensional, not binary: a language might have considerable written material but no spoken resources, or an active online community but nothing usable for computational processing. The resource tiers below track the availability and range of usable computational resources, not overall presence or cultural visibility online.
The deficit is not uniform across tasks. Written-text resources, though thin, exist for perhaps a hundred African languages. Speech resources — the recordings and transcriptions needed for speech recognition and synthesis — exist for far fewer, which is precisely inverted relative to need on a continent where a large share of languages are primarily oral.
| Languages | Resource level |
|---|---|
| ~30 | Benchmark datasets across multiple NLP tasks |
| ~100 | Some usable text corpus or parallel data |
| ~500 | Addressed by at least one language-identification or massively multilingual model |
| 1,500+ | No annotated resources of any kind |
Table 4 — Approximate number of African languages at each level of digital resource availability. Tiers are cumulative thresholds, not exclusive categories.
2. The exclusion is upstream of the data.
The data gap is a symptom of a research-capacity gap. In 2018, five of 2,695 author affiliations at the five major natural-language-processing conferences were African institutions. A field in which the continent holds a third of the world’s languages and under a fifth of one per cent of the research presence will not produce resources for those languages as a by-product of normal activity. It has to be deliberately funded.
That upstream exclusion matters for how interventions are designed. Commissioning datasets from outside the continent produces artifacts that are frequently unusable — wrong domain, wrong register, unreviewed transcription, no community consent, no maintenance path. The approaches that have worked have been participatory: built by researchers and speakers in the language communities concerned, with the resulting resources openly licensed.
3. What has been built
Progress since 2019 has been real and is worth stating precisely, both because it demonstrates feasibility and because it establishes the scale of what remains.
A grassroots participatory research community that has become one of the most influential networks in African NLP, producing machine translation across more than thirty languages and the benchmark datasets the field now relies on. Its 2026 LINGUA Africa program funds open datasets tied to real community use cases.
Named-entity and part-of-speech benchmarks covering around twenty typologically diverse African languages — the first resources of their kind, and still the reference point for evaluation.
News-domain machine translation data demonstrating that a few thousand high-quality translated sentences can meaningfully improve performance, which materially lowers the cost of adding a language.
Transformer language models adapted to African morphology and orthography, both outperforming general multilingual baselines on African tasks and establishing that architecture adaptation, not only data volume, matters.
Massively multilingual models and language-identification tools, the most ambitious of which addresses just over 500 African languages — the widest coverage achieved to date, and still a quarter of the continent.
Volunteer-contributed speech corpora and national initiatives have produced audio for a growing number of languages, though coverage remains far behind text and is concentrated in the largest languages.
4. Why it is an emergency
Three consequences follow from digital absence, and they compound.
Exclusion from services.As government, banking, health information and education move to digital delivery, they are delivered in the languages the systems support. A speaker whose language is unsupported is not offered a degraded service; they are offered service in someone else’s language, or none at all. The result is a direct transfer of disadvantage from linguistic minority status to material exclusion.
Accelerated shift. Where a language cannot be used in the digital domains where young people spend their attention, it becomes a language for home and elders. That status is a reliable precursor to interrupted transmission. Digital invisibility can contribute to language shift by limiting the contexts in which younger speakers can use their language, especially in education, communication, access to information, and digital media. This relationship may be reciprocal: languages already experiencing intergenerational shifts are less likely to develop digital resources, and the lack of such resources can further restrict the settings in which the language is spoken.
Foreclosure. Each generation of language technology is trained on the existing record. A language absent from the record when a model generation is trained is absent from everything built on that generation, and the gap widens rather than narrowing, because resource-rich languages accumulate resources faster. A closing window exists in which adding a language to the record is cheap; after it closes, retrofitting is far more expensive, and for languages that have lost their last fluent speakers, impossible.
The final consequence reframes the entire question of preservation. For a language with no recorded corpus, extinction is total: no route back exists, because nothing remains to revive from. For a language with a substantial audio and transcript record, even the loss of the last speaker is not final in the same way. Recording is not a substitute for revitalization, but for languages approaching the loss of fluent speakers it can provide an essential foundation for future documentation, teaching, research, and revitalization. The earlier a usable community-controlled record is created, the more options remain available to subsequent generations.
5. The dataset that does not exist
What is missing is specific and buildable: consented, openly licensed, speaker-verified audio with aligned transcription, at meaningful scale, across the full range of African languages rather than the best-resourced few dozen. Nothing about this undertaking is technically novel. The obstacles are coordination, funding, and the absence of any institution with a mandate for the whole continent rather than for a single language or country. Failing to create a comprehensive record of linguistic diversity will result in the next generation of language technology — and, consequently, a significant portion of educational, governmental, and commercial systems — being developed from an incomplete representation of the world’s languages. That absence will then be extremely difficult to correct, and the communities concerned will bear the cost of the correction rather than the institutions that permitted the omission.
Four situations, not four languages
Four structural conditions, each illustrated by several languages rather than one.
Single-language case studies invite the reader to treat each case as particular. The situations below are not: each describes a pattern affecting dozens or hundreds of African languages at once, illustrated by several languages rather than one, so an intervention designed for one member of a group works for the rest.
The four situations are not stages of a single progression. A language can be digitally thin without shifting, or shifting without being critically endangered. They are distinct conditions requiring distinct responses, and conflating them is one reason existing interventions have been poorly targeted.
| Vital, digitally thin | Shifting | Critical | Reversing | |
|---|---|---|---|---|
| Speakers | c. 400M | Tens of millions | Thousands to single digits | Tens of millions |
| Languages | 30–50 | Several hundred | 200–500 | A small group |
| Demographic risk | Low to none | Medium — two generations | High — under a decade | None — trajectory improving |
| What it needs | Volume, domain breadth, speech | Measurement, then documentation | Intensive specialist recording | Comparative study of method |
| Cost profile | High total, low per hour | Moderate, community-led | Low total, very high per hour | Research only |
Table 5 — The four situations compared.
VII.A — affects c. 30–50 languages · c. 400 million speakers
Large, healthy, and absent from the technology
These languages are in no demographic danger whatsoever. They have tens of millions of speakers each, unbroken intergenerational transmission, literary traditions, film and music industries, and in several cases official status. They are nonetheless close to invisible in the systems that now mediate access to information — which makes them the clearest possible demonstration that digital invisibility is not a function of speaker numbers.
| Language | Speakers | Note |
|---|---|---|
| Yoruba | c. 45M | Extensive literary and film tradition; tone-marking diacritics routinely dropped in digital text, degrading every downstream task. |
| Hausa | c. 80M | Major international broadcast presence; reaches the middle of the resource scale and still lacks robust speech data. |
| Amharic | c. 35M | Official language with its own script; Ge’ez rendering and input support remain inconsistent across platforms. |
| Oromo | c. 37M | Among the largest African languages by speakers and among the thinnest in digital resources relative to that size. |
| Igbo | c. 30M | Orthographic variation and diacritic loss fragment what corpora exist. |
| isiZulu, Shona, Wolof | 10–30M each | Regionally dominant, institutionally recognized, and outside the top resource tiers. |
Table 6 — Representative languages in the vital-but-digitally-thin group. Speaker figures are catalog estimates; see Section X.
Not one language in this group reaches the resource level of a mid-sized European language with a fraction of its speaker population. Where benchmarks exist, they cover a narrow set of tasks in a single domain, usually news text. Speech resources — the relevant modality for most of these speech communities — are thinner still. The gap is therefore not explained by market size or endangerment, because none of these languages is endangered.
Volume, breadth of domain, and speech. This group does not require rescue documentation; it requires ordinary contemporary language data at scale — conversational, broadcast, instructional and transactional material, openly licensed, with diacritics intact. It is also the group where investment returns fastest, because the speaker base is large enough to sustain the technologies built on the data.
VII.B — affects several hundred languages · the largest group by count
Losing ground to a lingua franca, one generation at a time
No standard measure yet endangers these languages. They have substantial speaker populations, often in the hundreds of thousands or millions, and are in active daily use. What is changing is the age profile of their speakers and the domains in which they are used. Children acquire them alongside a dominant regional language, use the dominant language in school and online, and raise their own children in it. Nothing in this process is visible in a speaker count until it is largely complete.
The evidence base fails this group most badly. Standard endangerment classifications rate these languages as safe or vulnerable on the strength of speaker counts that are frequently decades old, while the transmission data that would actually detect shift — what language children are acquiring, in which domains — has not been collected. A language can move from vigorous to moribund inside two generations, which is shorter than the interval since the last continental assessment.
Measurement first, then documentation. This group needs the reassessment that has not yet happened: age-stratified, domain-specific transmission data gathered at the community level. It is also the group where early documentation is cheapest and most effective, because fluent speakers are numerous and recording can be community-led rather than specialist-led. Acting here is preventive; acting later means acting on the next group instead.
VII.C — affects c. 200–500 languages, some with single-digit speaker numbers
Where the window closes within a decade
These languages have small, elderly speaker populations and no functioning transmission to children. Several have fewer than a dozen fluent speakers; some have one. For this group, the distinction between endangerment and digital invisibility collapses entirely, because when the last fluent speaker dies, whatever has been recorded is the entire remaining existence of the language. If nothing was recorded, the language does not become endangered — it becomes irrecoverable.
The documented extinction record for Africa lists at least 74 languages with no remaining native speakers and no spoken descendants, and that figure is certainly an undercount, since extinctions in unsurveyed areas go unrecorded. For most of those languages, the surviving record is a wordlist. The languages in this group are the ones for which that outcome is still avoidable, and the number of years in which it remains avoidable is small and known.
Speed and depth, immediately, with specialist support. This group requires intensive audio-first documentation with the remaining fluent speakers — connected speech, narrative, conversation, not word lists — and conventional crowdsourcing alone is unlikely to resolve it, as speaker communities are often very small, geographically dispersed, or mostly composed of elderly fluent speakers. These situations require intensive, locally coordinated documentation with specialized support. It is the smallest group by population and the most time-critical; delays in this group can lead to irreversible loss because the number of fluent speakers is critically low.
VII.D — a small group · the only evidence that trajectory is not fixed
Cases where the direction of travel changed
Decline is not the whole picture. A small number of African languages have measurably improved their position over the past three decades — through constitutional recognition, script adoption, broadcast presence, or deliberate community programs. These cases are worth examining closely because they indicate which levers actually move and roughly how much they cost.
The mechanisms in this group are not mysterious. Three recur: formal recognition that gives the language a role in the state; a usable, supported writing system, including keyboard and font availability; and a presence in media that young people actually consume. Two of the three are digital-infrastructure problems, consistent with the central argument here.
These cases suggest that positive attitudes alone are insufficient: durable change also requires meaningful domains in which speakers can use the language, including education, government, media, and digital communication.
Study, then replicate deliberately. These reversals occurred largely without coordination or documented method, and each is discussed anecdotally rather than analyzed. A structured comparative study of what was done, in what order, at what cost, would be one of the highest-value pieces of research available in this field — and is recommended here specifically.
The policy landscape
Inherited official-language arrangements, education, continental institutions, and consent.
Language policy in Africa is largely inherited. The official-language arrangements of most states were established at independence, in conditions that made a single administrative language attractive, and have been revised rarely since. The result is a continent where the languages of government are, in most cases, the languages of the former colonial power, and where the great majority of citizens conduct their public business in a language they did not learn at home.
1. Official status
Most African states recognize one or two official languages out of dozens spoken within their borders. A smaller number recognize a national language alongside the official one; a handful — South Africa, most prominently — recognize many. The distinction matters because official status determines whether a language appears in legislation, in courts, in public broadcasting, and in the documents citizens must read to access services. Languages without it are not banned; they are simply absent from every domain the state controls, which produces the same effect more slowly.
Where states have extended recognition, the effect has generally been positive and measurable — the Amazigh case discussed in Section VII.D is the clearest recent example. Recognition alone is not sufficient, however. Several constitutions name languages that receive no subsequent funding, materials, or institutional support, and nominal recognition without implementation has little observable effect on vitality.
2. Language in education
Numerous educational studies show that children tend to learn literacy and subject content more effectively when initial instruction is based on a language they already understand. However, the success of this approach depends on factors such as implementation, teacher preparation, teaching materials, and the overall educational context. The policy practice across most of the continent nonetheless involves a transition to an ex-colonial language within the first two to four years of primary schooling, frequently before literacy in any language is secure.
The usual justification is practical rather than pedagogical: no materials exist. That is accurate, and it is a data problem. Producing a full primary curriculum in a language requires a standardized orthography, a graded vocabulary, reference works, and a corpus large enough to draw examples from. For most African languages none of these exist, and the cost of producing them by conventional editorial means, language by language, is prohibitive. Here the digital-invisibility argument becomes concretely material: the corpus that would make mother-tongue education affordable is the same corpus missing from the technology.
Teacher supply compounds it. A teacher trained and examined in English or French, posted to a district whose language they do not speak, cannot deliver mother-tongue instruction regardless of policy. Language-of-instruction policy that does not address posting and training is not implementable.
3. Continental institutions
The African Union maintains a specialized institution for language policy, the African Academy of Languages (ACALAN), with a mandate that covers cross-border languages and the promotion of African languages for official use. The African Union has also elevated Kiswahili to working-language status and marked language-focused years and decades in its program cycle. ACALAN demonstrates that a continental institutional mandate is possible, but its mandate dwarfs its resources, raising doubts about its capacity to fund, staff, coordinate and implement. ACALAN’s budget and staffing are small relative to a remit covering two thousand languages across fifty-five member states, and its instruments are advisory.
Regional economic communities have taken on some language coordination, and national language boards and academies exist in several states with varying degrees of activity. What does not exist anywhere is an institution whose mandate is the continental language record — the data — as distinct from language promotion. No institution is charged with maintaining a comprehensive language record — a plausible reason no continent-wide vitality reassessment has happened since 2010.
4. Consent, ownership and data sovereignty
Language documentation in Africa has an extractive history. Material was collected from communities, deposited in institutions elsewhere, published under the collectors’ names, and in many cases never returned in usable form to the people who provided it. That history is directly relevant to present practice, because it is the reason some communities decline to participate in documentation and because the same pattern is now recurring in the training data.
The questions a documentation program must answer are specific. Who consented, to what use, and were they told that the material might train commercial systems? Who holds the recordings, and can the community obtain and re-use them? Can consent be withdrawn, and what happens to derived models if it is? Does the community receive anything when value is created downstream? Emerging Indigenous data-governance frameworks provide reasonable answers to these questions, and the African NLP research community has been notably active in insisting on data sovereignty and participatory methods as conditions of practice rather than optional extras.
Consent is not solely an ethical matter. Corpora collected without durable consent are legally fragile and cannot be safely published or built upon, which means extractive collection produces less usable data than participatory collection, not more.
5. What policy currently does not address
Three areas fall outside almost every national language policy on the continent. Digital infrastructure — orthographic standardization, Unicode coverage, fonts, keyboards, input methods, locale data — is treated as a technical matter for vendors rather than as public infrastructure, so no ministry owns it. Sign languages are recognized in a small number of constitutions and absent from most language policy instruments, and are excluded by construction from audio-based documentation programs. Urban vernaculars and contact languages, which are the actual daily medium for a large and growing share of African city-dwellers, are not recognized as languages at all in most policy frameworks and are therefore invisible to both preservation and education policy.
Recommendations
Specific enough to be accepted or rejected, ordered by tractability.
The recommendations below are stated with sufficient specificity to be accepted or rejected, rather than merely agreed with. They are ordered within each group by our assessment of tractability: the earlier items require less money and less coordination than the later ones.
IX.A — For governments and language bodies
- 01
Restore language questions to the census, with published methodology
The cheapest single improvement available. A language question with a documented coding scheme, asked consistently across successive censuses, would replace extrapolation with measurement for the majority of the continent’s speaker figures.
- 02
Publish orthographic standards as public infrastructure
Where a national language board exists, its orthographic decisions should be published in machine-readable form, with Unicode coverage confirmed and reference keyboard layouts released. Publishing the standard is a small, bounded, one-time cost per language with permanent downstream effect.
- 03
Treat input methods, fonts and locale data as a ministry responsibility
Assign ownership of digital language infrastructure to a named department. At present, no ministry in most states owns keyboard availability, font coverage, or locale support, so nobody is accountable when a national language cannot be typed on a standard device.
- 04
Delay the transition out of mother-tongue instruction, and fund the materials that make it possible
The pedagogical case is settled; the constraint is materials and teacher posting. Fund corpus and materials development as an explicit line item rather than treating the absence of materials as a reason to abandon the policy.
- 05
Recognize sign languages and urban contact languages in policy instruments
Both categories currently fall outside language policy altogether despite substantial user populations. Recognition is a precondition for their appearing in any education, broadcasting, or documentation programs.
- 06
Commission a continental vitality reassessment through ACALAN or an equivalent body
The evidence base is sixteen years old. A reassessment requires a mandate holder, a method, and funding on the order of a mid-sized statistical survey — not a research breakthrough. It is the highest-value item here, and the one least likely to happen without a specific decision.
IX.B — For funders and technologists
- 01
Fund audio-first, not text-first
For languages with no orthographic standard, a recording is unambiguous data and a transcript is a set of contested decisions. Audio-first collection defers to standardization rather than being blocked by it, and it aligns with the actual modality of most African speech communities.
- 02
Set breadth targets rather than depth targets
Coverage across many languages at a usable baseline is worth more than depth in a few. Current computational initiatives encompass several hundred African languages, but their coverage remains concentrated within a limited subset of the continent’s linguistic diversity.
- 03
Fund African institutions and community programs directly
Commissioned datasets produced outside the language community are frequently unusable — wrong domain, unreviewed transcription, no consent trail, no maintenance path. The participatory model has an established track record and should be the default rather than the exception.
- 04
Require open licensing and community data sovereignty as funding conditions
Corpora without durable consent are legally fragile and cannot safely be built upon — a quality requirement as much as an ethical one. Consent should be specific about downstream commercial and model-training use, and withdrawable.
- 05
Fund maintenance, not only creation
A dataset without a maintainer degrades: links rot, formats age, errors accumulate uncorrected. Grant structures that fund creation and not upkeep have produced a landscape of abandoned resources, and the marginal cost of maintaining an existing corpus is far below the cost of rebuilding it.
- 06
Fund the four situations separately
The groups in Section VIIneed opposite things — breadth and volume for the digitally thin, speed and depth for the critical, measurement for the shifting. Undifferentiated “African language preservation” funding tends to fund one while describing another.
IX.C — A research agenda
- 01
A comparative study of the reversal cases
The cases in Section VII.D succeeded without coordination and are discussed anecdotally. A structured account of what was done, in what order, at what cost, would be among the highest-value research available in this field.
- 02
Transmission data, not speaker counts
Age-stratified, domain-specific data on which language children are actually acquiring would detect shifts years before speaker counts change. No instrument currently exists for collecting transmission data at a continental scale.
- 03
A resource index maintained continuously
Digital resource availability is currently measured by occasional academic surveys that lag behind practice. A maintained per-language index, updated against public repositories, would make the deficit trackable rather than estimated only periodically.
- 04
Costing studies
Nobody has published a credible per-language cost for bringing an African language to a usable baseline of digital resources. Without that figure, funders cannot size the problem, and advocates cannot make a budgetary case.
- 05
The unknown-status languages
A residual set of languages appears in the catalogs on the strength of a single twentieth-century survey. Establishing whether they are still spoken is a bounded, finite piece of work that would materially improve every continental figure here.
Method, references and appendix
How each figure was derived, and how confident we are in it.
1. Method
This report synthesizes secondary sources; no new primary fieldwork was conducted. Findings are statements the underlying data or published research support directly; interpretations are conclusions drawn from patterns across sources; recommendations are normative proposals built on both. Every quantitative claim derives from one of the four principal reference catalogs, a national statistical return, or the peer-reviewed computational-linguistics literature, attributed in the text to its source and, where vintage matters, its year.
Where sources disagree, we report the range rather than a single figure; the disagreement is itself evidence about survey coverage. Where a widely cited figure traces only to an assessment more than a decade old, we say so at the point of use, rather than smoothing inconsistent figures into a single tidy dataset that would misrepresent the available confidence.
Confidence labeling
| Label | Definition | Applies to |
|---|---|---|
| Measured | Derived from a census or survey with published methodology, conducted within the last ten years. | A minority of large official languages. |
| Estimated | Catalog figure with a traceable source or an extrapolation from a survey more than 10 years old. | The majority of speaker figures in this report. |
| Stale | Widely cited figure traceable only to an assessment completed before 2011. | Most endangerment classifications, including UNESCO’s. |
| Unknown | No current assessment exists; the language is attested, but its present status is not established. | A substantial residual set, chiefly in Central Africa and South Sudan. |
Table 7 — Confidence labels used in the appendix and in quantitative claims throughout.
What is not measured
Not measured here: internet use, speaker attitudes, language prestige, intergenerational transmission, or the effectiveness of individual revitalization programs at continental scale. Nor does a lower digital-resource tier imply a language is less culturally or socially vital — the tier measures technological preparedness and representation, not sociolinguistic vitality.
Limitations
The principal catalogs were built for different purposes and share no common definitions, geographic boundaries, vitality criteria or update cycle. Combining them yields a comparative evidence base, not a harmonized census; their differences are preserved here rather than treating incompatible categories as equivalent.
Digital-resource estimates also age quickly: a language may gain a dataset, model or tool between a survey’s publication and this report’s release, and some resources are held privately or undocumented in public repositories. The figures here are a conservative snapshot of documented resources, not a complete inventory.
2. References
Adda, G., Stüker, S., Adda-Decker, M., et al. (2016). Breaking the unwritten language barrier: The BULB project. Procedia Computer Science, 81, 8–14. https://doi.org/10.1016/j.procs.2016.04.023
Adebara, I., & Abdul-Mageed, M. (2022). Towards Afrocentric NLP for African languages: Where we are and where we can go. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3814–3841. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.265
Adebara, I., Elmadany, A., & Abdul-Mageed, M. (2024). Cheetah: Natural language generation for 517 African languages. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12798–12823. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.691
Adebara, I., Elmadany, A., Abdul-Mageed, M., & Alcoba Inciarte, A. (2023). SERENGETI: Massively multilingual language models for Africa. Findings of the Association for Computational Linguistics: ACL 2023, 1498–1537. Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-acl.97
Adebara, I., Elmadany, A., Abdul-Mageed, M., & Inciarte, A. (2022). AfroLID: A neural language identification tool for African languages. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1958–1981. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.128
Adelani, D. I., Alabi, J. O., Fan, A., et al. (2022). A few thousand translations go a long way! Leveraging pre-trained models for African news translation. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3053–3070. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.naacl-main.223
Adelani, D. I., Neubig, G., Ruder, S., et al. (2022). MasakhaNER 2.0: Africa-centric transfer learning for named entity recognition. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4488–4508. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.298
African Union / African Academy of Languages (ACALAN). (n.d.). [Exact title of the institutional document or statute used].
Alabi, J. O., Adelani, D. I., Mosbach, M., & Klakow, D. (2022). Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning. Proceedings of the 29th International Conference on Computational Linguistics. Association for Computational Linguistics.
Batibo, H. M. (2005). Language decline and death in Africa: Causes, consequences and challenges. Multilingual Matters.
Caines, A. (2019). The geographic diversity of NLP conferences. Marginalia.
Childs, G. T. (2020). Language endangerment in Africa. Oxford Research Encyclopedia of Linguistics. Oxford University Press. https://doi.org/10.1093/acrefore/9780199384655.013.102
Dione, C. M. B., Adelani, D. I., Nabende, P., et al. (2023). MasakhaPOS: Part-of-speech tagging for typologically diverse African languages. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10883–10900. Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.609
Dryer, M. S., & Haspelmath, M. (Eds.). (2013). The World Atlas of Language Structures Online. Max Planck Institute for Evolutionary Anthropology.
Eberhard, D. M., Simons, G. F., & Fennig, C. D. (Eds.). (2026). Ethnologue: Languages of the world (29th ed.). SIL International.
Endangered Languages Project. (n.d.). Catalogue of Endangered Languages (ELCat). Retrieved September 15, 2026, from https://www.endangeredlanguages.com/
Hammarström, H., Forkel, R., Haspelmath, M., & Bank, S. (2026). Glottolog 5.3. Max Planck Institute for Evolutionary Anthropology. https://doi.org/10.5281/zenodo.18840935
Hedderich, M. A., Lange, L., Adel, H., Strötgen, J., & Klakow, D. (2021). A survey on recent approaches for natural language processing in low-resource scenarios. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics.
Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 6282–6293. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560
Martinus, L., & Abbott, J. Z. (2019). A focus on neural machine translation for African languages. arXiv. https://arxiv.org/abs/1906.05685
Moran, S., & McCloy, D. (Eds.). (2019). PHOIBLE 2.0. Max Planck Institute for the Science of Human History. https://doi.org/10.5281/zenodo.2626687
Masakhane. (2026). LINGUA Africa: Open call for inclusive AI language projects.
Nekoto, W., Marivate, V., Matsila, T., et al. (2020). Participatory research for low-resourced machine translation: A case study in African languages. Findings of the Association for Computational Linguistics: EMNLP 2020, 2144–2160. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.findings-emnlp.195
Ogueji, K., Zhu, Y., & Lin, J. (2021). Small data? No problem! Exploring the viability of pretrained multilingual language models for low-resourced languages. Proceedings of the 1st Workshop on Multilingual Representation Learning, 116–126. Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.mrl-1.11
Skirgård, H., Haynie, H. J., Blasi, D. E., et al. (2023). Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss. Science Advances, 9(16), eadg6175. https://doi.org/10.1126/sciadv.adg6175
UNESCO. (2002). Atlas of the world’s languages in danger of disappearing (2nd ed.). UNESCO Publishing.
UNESCO. (2010). Atlas of the world’s languages in danger (3rd ed.). UNESCO Publishing.
UNESCO. (n.d.). World Atlas of Languages. Retrieved September 15, 2026, from https://en.wal.unesco.org/
Language categorisation
How languages are classified here, and where the complete list lives.
| Dimension | Values | Source of the scheme |
|---|---|---|
| Family | Niger-Congo, Afro-Asiatic, Nilo-Saharan, Khoe-Kwadi / Tuu / Kx’a, Austronesian, Indo-European, contact languages, sign languages | Glottolog classification (Section IV.A) |
| Region | West, Central, East, Horn, Southern, North Africa | Six-region scheme (Section IV.B) |
| Vitality | Vital, Vulnerable, Endangered, Extinct, Unknown | UNESCO Atlas and ELCat, kept as separate unreconciled fields |
| Digital resource tier | T1 broad commercial support → T5 no annotated resource | Classification devised for this report |
| Confidence | MEASURED, ESTIMATED, STALE, UNKNOWN | Labeling scheme in Table 7 |
| Situation group | Vital/digitally thin, Shifting, Critical, Reversing | Case grouping in Section VII |
Table 8 — Categorization scheme applied throughout. See Section IV.A for family definitions, Section IV.B for the regional scheme, and Table 7 for confidence labeling.