Atilingo Research Report

State of African Languages: counting, vitality and digital emergency

A baseline assessment of approximately 2,000–2,200 living languages.

Published
September 2026
Licence
All rights reserved
Recommended citation
Atilingo Inc. (2026). The State of African Languages: counting, vitality, and the digital emergency across 2,000 languages.
Data sources
Glottolog · Ethnologue · UNESCO World Atlas of Languages · ELCat · national census agencies · published NLP resource surveys

About this report

Atilingo is a nonprofit technology organization. Its specific purpose is to operate exclusively for educational, scientific and charitable purposes: preserving, documenting and revitalizing Indigenous, endangered and underrepresented languages and the cultural knowledge and heritage associated with them; developing and supporting educational programs, digital tools and research that advance language learning, linguistic diversity and cultural preservation; and creating, developing and disseminating open, accessible resources and technologies that support the study, use and preservation of underrepresented languages. We work with speaker communities to record oral material, transcribe it with AI assistance and native-speaker review, and publish it as structured data that researchers, developers and the communities themselves can use.

This is the first in a planned annual series, setting out what is currently known — and, more importantly, what is not known — about the number, distribution, vitality and digital representation of African languages, as a baseline funders, ministries, researchers and technologists can cite from a common footing.

Evidence and method

Findings here draw on three kinds of evidence: established results from published research and language catalogs, estimates from incomplete or inconsistent data, and Atilingo’s own interpretations of patterns across sources. Where the evidence is uneven, we label the confidence — ranges, qualifications and “unknown” classifications are findings, not gaps.

The four principal reference works are Glottolog, Ethnologue, the UNESCO World Atlas of Languages, and the Catalog of Endangered Languages (ELCat), supplemented by national statistical returns and the published computational-linguistics literature. See Section X for method and the Appendix for the full enumeration.

Section I

Executive summary

Seven findings, the thesis, and the recommendations at a glance.

Africa holds roughly a third of the world’s languages and a small fraction of the world’s language data. The conventional framing of language loss as a preservation problem — communities losing speakers — is accurate but incomplete.

The more urgent emergency: the overwhelming majority of African languages are absent from the digital and computational record, and therefore absent from the technologies that will mediate education, government, commerce, and knowledge for the next century.

Seven findings

~2,000
Languages

No agreed count exists, and the disagreement is the finding

Published estimates for Africa range from about 1,250 to over 3,000. The catalogs diverge most precisely where survey coverage is thinnest, making the spread a usable proxy for how poorly documented a region is.

2010
Last update

The most-cited endangerment data is sixteen years old

UNESCO’s Atlas dataset was completed in 2010 and has not been substantively updated. Every widely quoted African endangerment figure predates the mobile-internet and AI era it is now used to describe.

428
Threatened

Several hundred languages are at risk, and up to a tenth may not survive the century

UNESCO records nearly 428 African languages as threatened; an independent survey identified over 500 as highly endangered, extinct or close to it. Both assessments are now more than fifteen years old.

<2%
Well resourced

Digital invisibility affects almost every African language, not only endangered ones

The majority of African languages have no annotated digital data of any kind. Fewer than two per cent meet the report’s higher digital-resource threshold — defined in the methodology as broad coverage across multiple computational tasks and usable text and/or speech resources — comparable to that available for a typical mid-sized European language.

5 of 2,695
Affiliations

The exclusion begins in the research pipeline

In 2018, five of 2,695 author affiliations at the five major NLP conferences were African institutions. Resources for African languages will not appear as a by-product of the field’s normal activity.

~890
In West Africa

Linguistic density and documentation are inversely related

West Africa holds roughly 890 living languages, more than any comparable region on earth, and is among the least surveyed. Central Africa shows the same pattern more acutely.

3
Omitted classes

Sign languages, creoles and urban vernaculars are systematically uncounted

African sign languages, contact languages such as Nigerian Pidgin with tens of millions of speakers, and fast-growing urban vernaculars sit almost entirely outside the reference catalogs.

Recommendations at a glance

For governments & language bodies
  • Fund a continental vitality reassessment; the current evidence base is sixteen years old.
  • Add or restore language questions to national census instruments, with published methodology.
  • Treat orthographic standardization and keyboard/input support as digital infrastructure rather than cultural policy.
  • Recognize mother-tongue instruction as a data problem as much as a pedagogical one — materials require corpora.
For funders & technologists
  • Fund audio-first, consented, openly licensed corpora rather than text-only or extractive collection.
  • Fund African research institutions and participatory community programs directly, not commissioned deliverables.
  • Set breadth targets, not depth targets: coverage across many languages beats depth in a few.
  • Require open licensing and community data sovereignty as a condition of funding, and fund maintenance, not only creation.
Section II

Introduction

Why this report exists, what it argues, and what it deliberately does not claim.

1. Why this report, and why now

Two developments make 2026 a reasonable moment to restate the position of Africa’s languages. The first: the most widely cited endangerment data for the continent has aged badly. UNESCO’s Atlas of the World’s Languages in Danger dataset was completed in 2010 and has not been substantively updated since, so the figures press coverage, funding proposals and policy documents reproduce describe a situation more than fifteen years old.

The second: a new and largely unmeasured form of language loss has emerged in that same interval. Between 2010 and 2026, the primary interfaces for information, government services, commerce, and education shifted decisively toward systems built on large text and speech corpora. Those corpora do not merely underserve the languages absent from them; they cannot see those languages at all. A language with two million healthy speakers and no digital footprint is, from the perspective of a search engine, a translation system or a speech interface, indistinguishable from a language with no speakers at all.

How many African languages there are, how they are faring, and to what extent those languages exist in the digital record are inseparable questions. The first is the traditional subject of language documentation. The second determines whether the answer to the first will still be true in fifty years.

2. The argument

Our thesis posits that the emergency facing African languages has taken on a second, increasingly significant dimension: digital invisibility. While the loss of speakers remains the most immediate threat to the survival of critically endangered languages, the lack of digital presence now affects a much broader range of languages, including those with substantial and active speaker communities. Consequently, the central question has evolved; it is no longer solely about whether a language has speakers, but rather whether it is represented in the digital infrastructures that mediate knowledge, education, public services, and communication. This is not an argument that endangerment does not matter. Several hundred African languages have small, aging speaker populations and will not survive the century without intervention, and the loss of each is irreversible. It argues instead about where the leverage now lies and how large the affected population is.

Classical endangerment affects a minority of African languages — those with the smallest and most isolated speaker communities. Digital invisibility affects almost all of them, including the largest and healthiest. Hausa, Yoruba, Amharic, Oromo, Igbo and Zulu together account for hundreds of millions of speakers, and all of them remain poorly represented in the systems that now mediate access to information. For the several hundred languages with fewer than ten thousand speakers, digital invisibility and endangerment are the same problem, because a language with no recorded material has no route to revitalization once transmission stops.

In practice, priorities must reorder. Documentation has historically prioritized urgency: record the most endangered languages first. That logic remains sound for the critical cases, but it produces an archive optimized for scholarship rather than for use. An approach organized around usability — machine-readable, openly licensed, audio-anchored data across the full range of languages — serves both the critical cases and the large, healthy languages simultaneously, and it is the only approach that scales to 2,000 languages.

3. Limits

No new primary fieldwork underlies this synthesis of existing catalogs, statistical returns, and published literature, and its accuracy is bounded accordingly. No single authoritative count of African languages exists, so none is offered here; the internal linguistic structure of the languages discussed goes unassessed. And technology does not substitute for intergenerational transmission: a language survives because children speak it, not because it has a dataset.

We also note our own position: Atilingo builds language datasets, and a report arguing that language datasets are urgently needed is not disinterested. We manage this by grounding every claim in third-party sources, confining our own work to a single short section, and publishing the underlying data openly so the conclusions can be checked against it.

Section III

How many languages does Africa have?

Four catalogs, four answers, and what the disagreement between them measures.

No agreed answer exists. Published estimates for the number of living languages in Africa range from about 1,250 to over 3,000, and the range is not the result of careless scholarship. It reflects genuine, unresolved disagreement about what counts as a distinct language, compounded by uneven survey coverage across the continent.

1. What the four principal catalogs say

The figures below are the continental totals implied by the four reference works drawn on here. They are not directly comparable, and the reasons they are not comparable are more informative than the numbers themselves.

Reference workContinental totalUnit countedAccess
Glottolog~2,300–2,400Genealogically distinguished languoids, split where published literature supports itOpen, CC BY, fully downloadable
Ethnologue~2,140–2,150Living languages; functional criterion rooted in translation need; ISO 639-3 registrarSubscription; licensing restricts republication
UNESCO WAL428 listedOnly languages assessed as endangered; silent on non-endangered languagesOpen; core dataset completed 2010
ELCatPartial coverageEndangered languages with structured vitality evidence and sourcesOpen, community-updatable

Table 1 — Continental language totals by reference work. Totals are approximate and change between editions; see Section X for edition-specific figures and retrieval dates.

2. Why the numbers differ

The language/dialect boundary is partly political. Whether two neighboring varieties are counted as one language or two frequently depends on whether their speakers are treated as one people or two. States, churches, publishers and census authorities make that determination, not linguists. Splitting a dialect continuum into separate entries raises the count; merging it lowers the count. Neither operation is neutral, and both have been used to political ends — to inflate the apparent unity of a nation, or to fragment a minority.

The catalogs have different purposes.Ethnologue’s lineage is in Bible translation, and its unit of interest has historically been the community requiring a separate translation — a functional criterion. Glottolog is built for genealogical classification and bibliographic reference, and it splits where the published literature supports a split. UNESCO’s Atlas exists to flag endangerment, so it enumerates only languages assessed as at risk and is silent on the rest.

Coverage is geographically uneven. The Bantu-speaking areas of Central Africa, parts of South Sudan, Chad and the Ethiopian periphery contain languages that have never been the subject of a published grammar or wordlist. Their existence is attested, sometimes only by a single survey visit decades ago, but their boundaries and speaker numbers are effectively unknown.

Systematic exclusions. Three categories are routinely under-counted. African sign languages — of which several dozen are attested, some with substantial user communities — are absent from most continental figures. Widely spoken contact languages and creoles are inconsistently treated: Nigerian Pidgin, with tens of millions of speakers, is frequently omitted from national language counts altogether. And varieties of Arabic and Amazigh are grouped and split inconsistently between sources, moving the North African total by dozens.

3. The figure used here

Where a single number is required: approximately 2,000 to 2,200 living languages, the band supported by the two catalogs with full continental coverage and open or documented methodology. The true figure depends on decisions that are not purely linguistic; per-language data is in the Appendix for readers who want to apply their own splitting criteria.

By this estimate, Africa holds approximately one-third of the world’s living languages — the most linguistically diverse continental region by number of languages.

Section IV

The linguistic landscape

Families, regions, writing systems, and the categories the count usually omits.

Africa’s languages are conventionally assigned to four large groupings, a scheme proposed in the 1960s that has not survived subsequent scrutiny intact. Two of the four are now seriously contested. The classification below reflects current consensus where it exists. It marks the disputes where it does not, because the dispute has practical consequences: a language whose family placement is unknown is typically also a language whose grammar is unpublished.

1. Families

Niger-Congo

~1,400–1,500 languages · c. 600 million speakers

The largest language family on earth by number of languages, spanning West, Central, East and Southern Africa. Its Bantu branch alone accounts for around 500 languages, including Swahili, Zulu, Shona and Kikuyu; its West African branches include Yoruba, Igbo, Akan and Fula.

Status of the classification: The family is widely accepted, but internal structure is unsettled, and the coherence of some proposed branches remains disputed. Many individual Bantu languages have no published grammar.

Afro-Asiatic

~375 languages · c. 500 million speakers

Spans North Africa, the Sahel and the Horn, and extends into the Middle East. Includes Arabic varieties, Amharic, Tigrinya, Somali, Oromo, Hausa and the Amazigh (Berber) languages.

Status of the classification: Well established as a family. Disputes concern the treatment of Arabic and Amazigh varieties — whether to count them as single languages or as dozens — which moves the continental total substantially.

Nilo-Saharan

~200 languages · c. 60 million speakers

A geographically dispersed grouping across the Sahel, Nile valley and Great Lakes, including Luo, Dinka, Nuer, Kanuri, Maasai and the Nubian languages.

Status of the classification: Contested. Many specialists regard Nilo-Saharan as a convenience grouping of several unrelated families rather than a demonstrated genetic unit. Nilo-Saharan is the least securely classified major grouping on the continent.

Khoe-Kwadi, Tuu, Kx’a

~30 languages · under 500,000 speakers

Southern African languages formerly bundled as “Khoisan”, including N|uu, Ju|’hoan, Khoekhoe and Naro. Distinguished by exceptionally large phonemic inventories, including extensive click systems.

Status of the classification:The single “Khoisan” family is now generally rejected; it is instead treated as three unrelated families plus isolates. They contain the most critically endangered languages in Africa, several of which have fewer than a dozen fluent speakers.

Austronesian and Indo-European

~10 languages · c. 40 million speakers

Malagasy, an Austronesian language spoken across Madagascar by some 25 million people, and Afrikaans, an Indo-European language with roots in Dutch, alongside the colonial languages that remain official in most African states.

Status of the classification:Uncontroversial. Included here because these languages are frequently omitted from “African languages” counts despite being first languages for tens of millions of Africans.

Contact languages, creoles and sign languages

Several dozen attested · tens of millions of users

Nigerian Pidgin, Krio, Cape Verdean Creole, Sango, Sheng and Camfranglais, alongside several dozen African sign languages from nationally recognized to village-scale.

Status of the classification: Inconsistently cataloged or absent entirely. It is the fastest-growing and least documented category of African language, and the one where the gap between speaker numbers and digital presence is widest.

2. Regional profiles

Linguistic density is not evenly spread. West Africa alone accounts for close to half of the continent’s languages, and Nigeria alone for around a quarter of West Africa’s total. The figures below are indicative midpoint estimates; per-country details are in the Appendix.

RegionLanguagesPrincipal familiesDominant pressures
West Africa~890Niger-Congo, Afro-Asiatic, Nilo-SaharanUrban lingua francas; English/French in education
Central Africa~680Niger-Congo (Bantu), Nilo-SaharanWeakest documentation; spread of French, Lingala and Sango
East Africa~440Niger-Congo (Bantu), Nilo-Saharan, Afro-AsiaticSwahili as regional lingua franca; English in schooling
Horn of Africa~180Afro-Asiatic, Nilo-SaharanAmharic and Somali dominance; conflict displacement
Southern Africa~130Niger-Congo (Bantu), Khoe-Kwadi, Tuu, Kx’aMost critical endangerment cases; English/Afrikaans shift
North Africa~60Afro-Asiatic (Arabic, Amazigh)Arabic diglossia; state policy toward Amazigh varieties

Table 2 — Regional distribution. Counts are mid-point estimates across catalogs and should be read as orders of magnitude, not precise totals.

3. Writing systems

Most African languages that are written at all use a Latin-based orthography, frequently one devised by missionaries or colonial administrations and never formally standardized. A minority use Ge’ez script (Amharic, Tigrinya, Tigre), Arabic script (including the Ajami traditions used for Hausa, Wolof, Fula and Swahili), or one of the indigenous scripts developed in the nineteenth and twentieth centuries — Vai, N’Ko, Tifinagh, Osmanya, Adlam, Mwangwego and others.

Orthographic instability has a direct digital cost. Where a language has no agreed spelling standard, text corpora fragment across competing conventions and every downstream computational task degrades. Where a language uses tone marking that keyboards do not readily produce, writers omit the diacritics, and the resulting text is ambiguous in ways that are difficult to recover automatically. A substantial share of the digital deficit documented in Section VI is not a shortage of speakers willing to write but an absence of the standards and input methods that would make their writing usable.

Orthographic instability is also the strongest argument for treating audio as the primary medium. For a language with no orthographic standard, a recording is unambiguous data; a transcript is a set of decisions. Audio-first documentation defers the question of standardization rather than being blocked by it.

4. Categories the count usually omits

Sign languages. Several dozen African sign languages are attested, ranging from national languages with formal recognition to village sign languages with a few dozen users. They are absent from most continental counts, almost absent from digital resources, and, by construction, excluded from any documentation program built on audio.

Contact languages and urban vernaculars. Nigerian Pidgin, Sheng, Camfranglais, Tsotsitaal and similar varieties have very large and growing user populations, are the dominant medium of daily life for many urban Africans, and sit almost entirely outside the reference catalogs. They are simultaneously the fastest-growing and least documented category on the continent.

Languages of recent record only.A residual set of languages appears in the catalogs on the strength of a single twentieth-century survey and has not been verified since. For these, the honest status is not “endangered” or “vital” but unknown — and unknown status is itself a finding that current classification schemes have no way to express.

Section V

Vitality, endangerment and extinction

What the assessments record, what drives shift, and why the evidence base has aged badly.

Africa’s endangerment profile differs from that of the Americas, Australia or Siberia. Proportionally fewer African languages are moribund, because a large number of mid-sized languages remain in vigorous daily use across multiple generations. Some read that as reassurance. It should not be: the mechanisms that produced rapid language loss elsewhere are present and intensifying, and the data available to monitor them is worse than for any other continent.

1. What the assessments record

The principal published assessments are broadly consistent in magnitude while differing in method and vintage.

AssessmentVintageWhat it records
UNESCO AtlasDataset completed 2010Records nearly 428 African languages as threatened across four tiers and warns that up to 10 per cent of African languages may vanish within a century. Not substantively updated since completion.
UNESCO Atlas (earlier edition)c. 2002Of the roughly 1,400 African languages then cataloged, between 500 and 600 were assessed as in decline and 250 as under immediate threat. The same edition described Africa as linguistically the least-known continent.
Batibo continental survey2005A country-by-country survey identified 509 African languages as highly endangered, extinct or on the brink of extinction — broadly corroborating the UNESCO magnitude by an independent method.
Documented extinctionsCumulativeAt least 74 languages that diverged from their parent language in Africa are recorded as having no remaining native speakers and no spoken descendants. The actual figure is likely higher, because extinctions in poorly surveyed areas may go unrecorded.

Table 3 — Principal published vitality assessments for Africa.

2. Drivers of shift

Education in ex-colonial languages. In most African states, the language of instruction shifts to English, French, Portuguese or Arabic within the first few years of primary school, marking the home language as inadequate for formal knowledge at the exact age when children are calibrating which language carries status.

Urbanization and regional lingua francas. Rural-to-urban migration places speakers of many languages in shared space, where a lingua franca — Swahili, Hausa, Wolof, Lingala, Amharic, Arabic — becomes the medium of daily life. In some communities, such processes can produce substantial intergenerational shift within two generations, although the pace varies considerably by language, region, domain of use, and community.

Absence of official status. A language with no role in government, courts, media or commerce offers its speakers no instrumental reason to maintain it. Official-language policy in most African states recognizes one or two languages out of dozens, and the unrecognized majority carry the full cost of that choice.

Small and dispersed speaker populations. Several hundred African languages have fewer than 10,000 speakers, and for many the community is not geographically concentrated. Below a certain population and density, ordinary demographic pressure is sufficient to end transmission without any specific hostile policy.

Digital absence. Treated at length in Section VI. Since the last continental vitality assessment, digital absence has moved from a marginal factor to a primary one, and it is the only driver on this list that current assessment instruments do not measure at all.

3. A note on measuring vitality

Speaker counts are the weakest link in the chain. Most published figures for African languages derive from national censuses, and censuses ask about language inconsistently, infrequently, or not at all. Some states do not collect language data because the results would be politically inconvenient. Others report figures unchanged across successive editions, which cannot be accurate for a continent whose population is growing rapidly. A significant number of the speaker estimates in circulation are extrapolations from decades-old surveys, carried forward and re-cited until their provenance is lost.

A speaker count alone also captures vitality poorly. A language with a million speakers, none of them under thirty, is in more danger than a language with five thousand speakers being raised as children’s first language. The standard assessment scales attempt to capture that dynamic through transmission-based criteria, but applying them requires community-level observation that has not been carried out for most African languages. For a substantial proportion of the continent’s languages, no current vitality assessment exists — and that omission is itself the honest assessment.

Section VI — core chapter

Digital invisibility

Absence from the data is absence from the century.

1. The shape of the deficit

African languages are near-absent from the large web-crawled corpora on which contemporary language technology is built. Their representation in Wikipedia, and in web crawls generally, is a small fraction of a per cent. Because those crawls are the raw material for translation systems, speech recognition and large language models, the absence propagates through every downstream system. Computational-linguistics literature describes the resulting condition as a digital language divide, in which standard processing pipelines simply fail to generalize to African linguistic contexts.

The widely used resource taxonomy in that literature places languages on a six-level scale based on the availability of labeled and unlabelled data. The great majority of African languages sit at the bottom two levels — the categories informally named the left-behinds and the scraping-bys — with no annotated data of any kind. A handful, notably Swahili, Hausa, Amharic and Afrikaans, reach the middle of the scale.

Digital-resource status is multidimensional, not binary: a language might have considerable written material but no spoken resources, or an active online community but nothing usable for computational processing. The resource tiers below track the availability and range of usable computational resources, not overall presence or cultural visibility online.

The deficit is not uniform across tasks. Written-text resources, though thin, exist for perhaps a hundred African languages. Speech resources — the recordings and transcriptions needed for speech recognition and synthesis — exist for far fewer, which is precisely inverted relative to need on a continent where a large share of languages are primarily oral.

LanguagesResource level
~30Benchmark datasets across multiple NLP tasks
~100Some usable text corpus or parallel data
~500Addressed by at least one language-identification or massively multilingual model
1,500+No annotated resources of any kind

Table 4 — Approximate number of African languages at each level of digital resource availability. Tiers are cumulative thresholds, not exclusive categories.

2. The exclusion is upstream of the data.

The data gap is a symptom of a research-capacity gap. In 2018, five of 2,695 author affiliations at the five major natural-language-processing conferences were African institutions. A field in which the continent holds a third of the world’s languages and under a fifth of one per cent of the research presence will not produce resources for those languages as a by-product of normal activity. It has to be deliberately funded.

That upstream exclusion matters for how interventions are designed. Commissioning datasets from outside the continent produces artifacts that are frequently unusable — wrong domain, wrong register, unreviewed transcription, no community consent, no maintenance path. The approaches that have worked have been participatory: built by researchers and speakers in the language communities concerned, with the resulting resources openly licensed.

3. What has been built

Progress since 2019 has been real and is worth stating precisely, both because it demonstrates feasibility and because it establishes the scale of what remains.

Masakhane

A grassroots participatory research community that has become one of the most influential networks in African NLP, producing machine translation across more than thirty languages and the benchmark datasets the field now relies on. Its 2026 LINGUA Africa program funds open datasets tied to real community use cases.

MasakhaNER / MasakhaPOS

Named-entity and part-of-speech benchmarks covering around twenty typologically diverse African languages — the first resources of their kind, and still the reference point for evaluation.

MAFAND-MT

News-domain machine translation data demonstrating that a few thousand high-quality translated sentences can meaningfully improve performance, which materially lowers the cost of adding a language.

AfriBERTa, AfroXLMR

Transformer language models adapted to African morphology and orthography, both outperforming general multilingual baselines on African tasks and establishing that architecture adaptation, not only data volume, matters.

SERENGETI, Cheetah, AfroLID

Massively multilingual models and language-identification tools, the most ambitious of which addresses just over 500 African languages — the widest coverage achieved to date, and still a quarter of the continent.

Community speech collection

Volunteer-contributed speech corpora and national initiatives have produced audio for a growing number of languages, though coverage remains far behind text and is concentrated in the largest languages.

4. Why it is an emergency

Three consequences follow from digital absence, and they compound.

Exclusion from services.As government, banking, health information and education move to digital delivery, they are delivered in the languages the systems support. A speaker whose language is unsupported is not offered a degraded service; they are offered service in someone else’s language, or none at all. The result is a direct transfer of disadvantage from linguistic minority status to material exclusion.

Accelerated shift. Where a language cannot be used in the digital domains where young people spend their attention, it becomes a language for home and elders. That status is a reliable precursor to interrupted transmission. Digital invisibility can contribute to language shift by limiting the contexts in which younger speakers can use their language, especially in education, communication, access to information, and digital media. This relationship may be reciprocal: languages already experiencing intergenerational shifts are less likely to develop digital resources, and the lack of such resources can further restrict the settings in which the language is spoken.

Foreclosure. Each generation of language technology is trained on the existing record. A language absent from the record when a model generation is trained is absent from everything built on that generation, and the gap widens rather than narrowing, because resource-rich languages accumulate resources faster. A closing window exists in which adding a language to the record is cheap; after it closes, retrofitting is far more expensive, and for languages that have lost their last fluent speakers, impossible.

The final consequence reframes the entire question of preservation. For a language with no recorded corpus, extinction is total: no route back exists, because nothing remains to revive from. For a language with a substantial audio and transcript record, even the loss of the last speaker is not final in the same way. Recording is not a substitute for revitalization, but for languages approaching the loss of fluent speakers it can provide an essential foundation for future documentation, teaching, research, and revitalization. The earlier a usable community-controlled record is created, the more options remain available to subsequent generations.

5. The dataset that does not exist

What is missing is specific and buildable: consented, openly licensed, speaker-verified audio with aligned transcription, at meaningful scale, across the full range of African languages rather than the best-resourced few dozen. Nothing about this undertaking is technically novel. The obstacles are coordination, funding, and the absence of any institution with a mandate for the whole continent rather than for a single language or country. Failing to create a comprehensive record of linguistic diversity will result in the next generation of language technology — and, consequently, a significant portion of educational, governmental, and commercial systems — being developed from an incomplete representation of the world’s languages. That absence will then be extremely difficult to correct, and the communities concerned will bear the cost of the correction rather than the institutions that permitted the omission.

Section VII

Four situations, not four languages

Four structural conditions, each illustrated by several languages rather than one.

Single-language case studies invite the reader to treat each case as particular. The situations below are not: each describes a pattern affecting dozens or hundreds of African languages at once, illustrated by several languages rather than one, so an intervention designed for one member of a group works for the rest.

The four situations are not stages of a single progression. A language can be digitally thin without shifting, or shifting without being critically endangered. They are distinct conditions requiring distinct responses, and conflating them is one reason existing interventions have been poorly targeted.

Vital, digitally thinShiftingCriticalReversing
Speakersc. 400MTens of millionsThousands to single digitsTens of millions
Languages30–50Several hundred200–500A small group
Demographic riskLow to noneMedium — two generationsHigh — under a decadeNone — trajectory improving
What it needsVolume, domain breadth, speechMeasurement, then documentationIntensive specialist recordingComparative study of method
Cost profileHigh total, low per hourModerate, community-ledLow total, very high per hourResearch only

Table 5 — The four situations compared.

VII.A — affects c. 30–50 languages · c. 400 million speakers

Large, healthy, and absent from the technology

These languages are in no demographic danger whatsoever. They have tens of millions of speakers each, unbroken intergenerational transmission, literary traditions, film and music industries, and in several cases official status. They are nonetheless close to invisible in the systems that now mediate access to information — which makes them the clearest possible demonstration that digital invisibility is not a function of speaker numbers.

LanguageSpeakersNote
Yorubac. 45MExtensive literary and film tradition; tone-marking diacritics routinely dropped in digital text, degrading every downstream task.
Hausac. 80MMajor international broadcast presence; reaches the middle of the resource scale and still lacks robust speech data.
Amharicc. 35MOfficial language with its own script; Ge’ez rendering and input support remain inconsistent across platforms.
Oromoc. 37MAmong the largest African languages by speakers and among the thinnest in digital resources relative to that size.
Igboc. 30MOrthographic variation and diacritic loss fragment what corpora exist.
isiZulu, Shona, Wolof10–30M eachRegionally dominant, institutionally recognized, and outside the top resource tiers.

Table 6 — Representative languages in the vital-but-digitally-thin group. Speaker figures are catalog estimates; see Section X.

Not one language in this group reaches the resource level of a mid-sized European language with a fraction of its speaker population. Where benchmarks exist, they cover a narrow set of tasks in a single domain, usually news text. Speech resources — the relevant modality for most of these speech communities — are thinner still. The gap is therefore not explained by market size or endangerment, because none of these languages is endangered.

What this group needs

Volume, breadth of domain, and speech. This group does not require rescue documentation; it requires ordinary contemporary language data at scale — conversational, broadcast, instructional and transactional material, openly licensed, with diacritics intact. It is also the group where investment returns fastest, because the speaker base is large enough to sustain the technologies built on the data.

VII.B — affects several hundred languages · the largest group by count

Losing ground to a lingua franca, one generation at a time

No standard measure yet endangers these languages. They have substantial speaker populations, often in the hundreds of thousands or millions, and are in active daily use. What is changing is the age profile of their speakers and the domains in which they are used. Children acquire them alongside a dominant regional language, use the dominant language in school and online, and raise their own children in it. Nothing in this process is visible in a speaker count until it is largely complete.

The evidence base fails this group most badly. Standard endangerment classifications rate these languages as safe or vulnerable on the strength of speaker counts that are frequently decades old, while the transmission data that would actually detect shift — what language children are acquiring, in which domains — has not been collected. A language can move from vigorous to moribund inside two generations, which is shorter than the interval since the last continental assessment.

What this group needs

Measurement first, then documentation. This group needs the reassessment that has not yet happened: age-stratified, domain-specific transmission data gathered at the community level. It is also the group where early documentation is cheapest and most effective, because fluent speakers are numerous and recording can be community-led rather than specialist-led. Acting here is preventive; acting later means acting on the next group instead.

VII.C — affects c. 200–500 languages, some with single-digit speaker numbers

Where the window closes within a decade

These languages have small, elderly speaker populations and no functioning transmission to children. Several have fewer than a dozen fluent speakers; some have one. For this group, the distinction between endangerment and digital invisibility collapses entirely, because when the last fluent speaker dies, whatever has been recorded is the entire remaining existence of the language. If nothing was recorded, the language does not become endangered — it becomes irrecoverable.

The documented extinction record for Africa lists at least 74 languages with no remaining native speakers and no spoken descendants, and that figure is certainly an undercount, since extinctions in unsurveyed areas go unrecorded. For most of those languages, the surviving record is a wordlist. The languages in this group are the ones for which that outcome is still avoidable, and the number of years in which it remains avoidable is small and known.

What this group needs

Speed and depth, immediately, with specialist support. This group requires intensive audio-first documentation with the remaining fluent speakers — connected speech, narrative, conversation, not word lists — and conventional crowdsourcing alone is unlikely to resolve it, as speaker communities are often very small, geographically dispersed, or mostly composed of elderly fluent speakers. These situations require intensive, locally coordinated documentation with specialized support. It is the smallest group by population and the most time-critical; delays in this group can lead to irreversible loss because the number of fluent speakers is critically low.

VII.D — a small group · the only evidence that trajectory is not fixed

Cases where the direction of travel changed

Decline is not the whole picture. A small number of African languages have measurably improved their position over the past three decades — through constitutional recognition, script adoption, broadcast presence, or deliberate community programs. These cases are worth examining closely because they indicate which levers actually move and roughly how much they cost.

The mechanisms in this group are not mysterious. Three recur: formal recognition that gives the language a role in the state; a usable, supported writing system, including keyboard and font availability; and a presence in media that young people actually consume. Two of the three are digital-infrastructure problems, consistent with the central argument here.

These cases suggest that positive attitudes alone are insufficient: durable change also requires meaningful domains in which speakers can use the language, including education, government, media, and digital communication.

What this group needs

Study, then replicate deliberately. These reversals occurred largely without coordination or documented method, and each is discussed anecdotally rather than analyzed. A structured comparative study of what was done, in what order, at what cost, would be one of the highest-value pieces of research available in this field — and is recommended here specifically.

Section VIII

The policy landscape

Inherited official-language arrangements, education, continental institutions, and consent.

Language policy in Africa is largely inherited. The official-language arrangements of most states were established at independence, in conditions that made a single administrative language attractive, and have been revised rarely since. The result is a continent where the languages of government are, in most cases, the languages of the former colonial power, and where the great majority of citizens conduct their public business in a language they did not learn at home.

1. Official status

Most African states recognize one or two official languages out of dozens spoken within their borders. A smaller number recognize a national language alongside the official one; a handful — South Africa, most prominently — recognize many. The distinction matters because official status determines whether a language appears in legislation, in courts, in public broadcasting, and in the documents citizens must read to access services. Languages without it are not banned; they are simply absent from every domain the state controls, which produces the same effect more slowly.

Where states have extended recognition, the effect has generally been positive and measurable — the Amazigh case discussed in Section VII.D is the clearest recent example. Recognition alone is not sufficient, however. Several constitutions name languages that receive no subsequent funding, materials, or institutional support, and nominal recognition without implementation has little observable effect on vitality.

2. Language in education

Numerous educational studies show that children tend to learn literacy and subject content more effectively when initial instruction is based on a language they already understand. However, the success of this approach depends on factors such as implementation, teacher preparation, teaching materials, and the overall educational context. The policy practice across most of the continent nonetheless involves a transition to an ex-colonial language within the first two to four years of primary schooling, frequently before literacy in any language is secure.

The usual justification is practical rather than pedagogical: no materials exist. That is accurate, and it is a data problem. Producing a full primary curriculum in a language requires a standardized orthography, a graded vocabulary, reference works, and a corpus large enough to draw examples from. For most African languages none of these exist, and the cost of producing them by conventional editorial means, language by language, is prohibitive. Here the digital-invisibility argument becomes concretely material: the corpus that would make mother-tongue education affordable is the same corpus missing from the technology.

Teacher supply compounds it. A teacher trained and examined in English or French, posted to a district whose language they do not speak, cannot deliver mother-tongue instruction regardless of policy. Language-of-instruction policy that does not address posting and training is not implementable.

3. Continental institutions

The African Union maintains a specialized institution for language policy, the African Academy of Languages (ACALAN), with a mandate that covers cross-border languages and the promotion of African languages for official use. The African Union has also elevated Kiswahili to working-language status and marked language-focused years and decades in its program cycle. ACALAN demonstrates that a continental institutional mandate is possible, but its mandate dwarfs its resources, raising doubts about its capacity to fund, staff, coordinate and implement. ACALAN’s budget and staffing are small relative to a remit covering two thousand languages across fifty-five member states, and its instruments are advisory.

Regional economic communities have taken on some language coordination, and national language boards and academies exist in several states with varying degrees of activity. What does not exist anywhere is an institution whose mandate is the continental language record — the data — as distinct from language promotion. No institution is charged with maintaining a comprehensive language record — a plausible reason no continent-wide vitality reassessment has happened since 2010.

4. Consent, ownership and data sovereignty

Language documentation in Africa has an extractive history. Material was collected from communities, deposited in institutions elsewhere, published under the collectors’ names, and in many cases never returned in usable form to the people who provided it. That history is directly relevant to present practice, because it is the reason some communities decline to participate in documentation and because the same pattern is now recurring in the training data.

The questions a documentation program must answer are specific. Who consented, to what use, and were they told that the material might train commercial systems? Who holds the recordings, and can the community obtain and re-use them? Can consent be withdrawn, and what happens to derived models if it is? Does the community receive anything when value is created downstream? Emerging Indigenous data-governance frameworks provide reasonable answers to these questions, and the African NLP research community has been notably active in insisting on data sovereignty and participatory methods as conditions of practice rather than optional extras.

Consent is not solely an ethical matter. Corpora collected without durable consent are legally fragile and cannot be safely published or built upon, which means extractive collection produces less usable data than participatory collection, not more.

5. What policy currently does not address

Three areas fall outside almost every national language policy on the continent. Digital infrastructure — orthographic standardization, Unicode coverage, fonts, keyboards, input methods, locale data — is treated as a technical matter for vendors rather than as public infrastructure, so no ministry owns it. Sign languages are recognized in a small number of constitutions and absent from most language policy instruments, and are excluded by construction from audio-based documentation programs. Urban vernaculars and contact languages, which are the actual daily medium for a large and growing share of African city-dwellers, are not recognized as languages at all in most policy frameworks and are therefore invisible to both preservation and education policy.

Section IX

Recommendations

Specific enough to be accepted or rejected, ordered by tractability.

The recommendations below are stated with sufficient specificity to be accepted or rejected, rather than merely agreed with. They are ordered within each group by our assessment of tractability: the earlier items require less money and less coordination than the later ones.

IX.A — For governments and language bodies

  1. 01

    Restore language questions to the census, with published methodology

    The cheapest single improvement available. A language question with a documented coding scheme, asked consistently across successive censuses, would replace extrapolation with measurement for the majority of the continent’s speaker figures.

  2. 02

    Publish orthographic standards as public infrastructure

    Where a national language board exists, its orthographic decisions should be published in machine-readable form, with Unicode coverage confirmed and reference keyboard layouts released. Publishing the standard is a small, bounded, one-time cost per language with permanent downstream effect.

  3. 03

    Treat input methods, fonts and locale data as a ministry responsibility

    Assign ownership of digital language infrastructure to a named department. At present, no ministry in most states owns keyboard availability, font coverage, or locale support, so nobody is accountable when a national language cannot be typed on a standard device.

  4. 04

    Delay the transition out of mother-tongue instruction, and fund the materials that make it possible

    The pedagogical case is settled; the constraint is materials and teacher posting. Fund corpus and materials development as an explicit line item rather than treating the absence of materials as a reason to abandon the policy.

  5. 05

    Recognize sign languages and urban contact languages in policy instruments

    Both categories currently fall outside language policy altogether despite substantial user populations. Recognition is a precondition for their appearing in any education, broadcasting, or documentation programs.

  6. 06

    Commission a continental vitality reassessment through ACALAN or an equivalent body

    The evidence base is sixteen years old. A reassessment requires a mandate holder, a method, and funding on the order of a mid-sized statistical survey — not a research breakthrough. It is the highest-value item here, and the one least likely to happen without a specific decision.

IX.B — For funders and technologists

  1. 01

    Fund audio-first, not text-first

    For languages with no orthographic standard, a recording is unambiguous data and a transcript is a set of contested decisions. Audio-first collection defers to standardization rather than being blocked by it, and it aligns with the actual modality of most African speech communities.

  2. 02

    Set breadth targets rather than depth targets

    Coverage across many languages at a usable baseline is worth more than depth in a few. Current computational initiatives encompass several hundred African languages, but their coverage remains concentrated within a limited subset of the continent’s linguistic diversity.

  3. 03

    Fund African institutions and community programs directly

    Commissioned datasets produced outside the language community are frequently unusable — wrong domain, unreviewed transcription, no consent trail, no maintenance path. The participatory model has an established track record and should be the default rather than the exception.

  4. 04

    Require open licensing and community data sovereignty as funding conditions

    Corpora without durable consent are legally fragile and cannot safely be built upon — a quality requirement as much as an ethical one. Consent should be specific about downstream commercial and model-training use, and withdrawable.

  5. 05

    Fund maintenance, not only creation

    A dataset without a maintainer degrades: links rot, formats age, errors accumulate uncorrected. Grant structures that fund creation and not upkeep have produced a landscape of abandoned resources, and the marginal cost of maintaining an existing corpus is far below the cost of rebuilding it.

  6. 06

    Fund the four situations separately

    The groups in Section VIIneed opposite things — breadth and volume for the digitally thin, speed and depth for the critical, measurement for the shifting. Undifferentiated “African language preservation” funding tends to fund one while describing another.

IX.C — A research agenda

  1. 01

    A comparative study of the reversal cases

    The cases in Section VII.D succeeded without coordination and are discussed anecdotally. A structured account of what was done, in what order, at what cost, would be among the highest-value research available in this field.

  2. 02

    Transmission data, not speaker counts

    Age-stratified, domain-specific data on which language children are actually acquiring would detect shifts years before speaker counts change. No instrument currently exists for collecting transmission data at a continental scale.

  3. 03

    A resource index maintained continuously

    Digital resource availability is currently measured by occasional academic surveys that lag behind practice. A maintained per-language index, updated against public repositories, would make the deficit trackable rather than estimated only periodically.

  4. 04

    Costing studies

    Nobody has published a credible per-language cost for bringing an African language to a usable baseline of digital resources. Without that figure, funders cannot size the problem, and advocates cannot make a budgetary case.

  5. 05

    The unknown-status languages

    A residual set of languages appears in the catalogs on the strength of a single twentieth-century survey. Establishing whether they are still spoken is a bounded, finite piece of work that would materially improve every continental figure here.

Section X

Method, references and appendix

How each figure was derived, and how confident we are in it.

1. Method

This report synthesizes secondary sources; no new primary fieldwork was conducted. Findings are statements the underlying data or published research support directly; interpretations are conclusions drawn from patterns across sources; recommendations are normative proposals built on both. Every quantitative claim derives from one of the four principal reference catalogs, a national statistical return, or the peer-reviewed computational-linguistics literature, attributed in the text to its source and, where vintage matters, its year.

Where sources disagree, we report the range rather than a single figure; the disagreement is itself evidence about survey coverage. Where a widely cited figure traces only to an assessment more than a decade old, we say so at the point of use, rather than smoothing inconsistent figures into a single tidy dataset that would misrepresent the available confidence.

Confidence labeling

LabelDefinitionApplies to
MeasuredDerived from a census or survey with published methodology, conducted within the last ten years.A minority of large official languages.
EstimatedCatalog figure with a traceable source or an extrapolation from a survey more than 10 years old.The majority of speaker figures in this report.
StaleWidely cited figure traceable only to an assessment completed before 2011.Most endangerment classifications, including UNESCO’s.
UnknownNo current assessment exists; the language is attested, but its present status is not established.A substantial residual set, chiefly in Central Africa and South Sudan.

Table 7 — Confidence labels used in the appendix and in quantitative claims throughout.

What is not measured

Not measured here: internet use, speaker attitudes, language prestige, intergenerational transmission, or the effectiveness of individual revitalization programs at continental scale. Nor does a lower digital-resource tier imply a language is less culturally or socially vital — the tier measures technological preparedness and representation, not sociolinguistic vitality.

Limitations

The principal catalogs were built for different purposes and share no common definitions, geographic boundaries, vitality criteria or update cycle. Combining them yields a comparative evidence base, not a harmonized census; their differences are preserved here rather than treating incompatible categories as equivalent.

Digital-resource estimates also age quickly: a language may gain a dataset, model or tool between a survey’s publication and this report’s release, and some resources are held privately or undocumented in public repositories. The figures here are a conservative snapshot of documented resources, not a complete inventory.

2. References

Adda, G., Stüker, S., Adda-Decker, M., et al. (2016). Breaking the unwritten language barrier: The BULB project. Procedia Computer Science, 81, 8–14. https://doi.org/10.1016/j.procs.2016.04.023

Adebara, I., & Abdul-Mageed, M. (2022). Towards Afrocentric NLP for African languages: Where we are and where we can go. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3814–3841. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.265

Adebara, I., Elmadany, A., & Abdul-Mageed, M. (2024). Cheetah: Natural language generation for 517 African languages. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12798–12823. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.691

Adebara, I., Elmadany, A., Abdul-Mageed, M., & Alcoba Inciarte, A. (2023). SERENGETI: Massively multilingual language models for Africa. Findings of the Association for Computational Linguistics: ACL 2023, 1498–1537. Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-acl.97

Adebara, I., Elmadany, A., Abdul-Mageed, M., & Inciarte, A. (2022). AfroLID: A neural language identification tool for African languages. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1958–1981. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.128

Adelani, D. I., Alabi, J. O., Fan, A., et al. (2022). A few thousand translations go a long way! Leveraging pre-trained models for African news translation. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3053–3070. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.naacl-main.223

Adelani, D. I., Neubig, G., Ruder, S., et al. (2022). MasakhaNER 2.0: Africa-centric transfer learning for named entity recognition. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4488–4508. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.298

African Union / African Academy of Languages (ACALAN). (n.d.). [Exact title of the institutional document or statute used].

Alabi, J. O., Adelani, D. I., Mosbach, M., & Klakow, D. (2022). Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning. Proceedings of the 29th International Conference on Computational Linguistics. Association for Computational Linguistics.

Batibo, H. M. (2005). Language decline and death in Africa: Causes, consequences and challenges. Multilingual Matters.

Caines, A. (2019). The geographic diversity of NLP conferences. Marginalia.

Childs, G. T. (2020). Language endangerment in Africa. Oxford Research Encyclopedia of Linguistics. Oxford University Press. https://doi.org/10.1093/acrefore/9780199384655.013.102

Dione, C. M. B., Adelani, D. I., Nabende, P., et al. (2023). MasakhaPOS: Part-of-speech tagging for typologically diverse African languages. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10883–10900. Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.609

Dryer, M. S., & Haspelmath, M. (Eds.). (2013). The World Atlas of Language Structures Online. Max Planck Institute for Evolutionary Anthropology.

Eberhard, D. M., Simons, G. F., & Fennig, C. D. (Eds.). (2026). Ethnologue: Languages of the world (29th ed.). SIL International.

Endangered Languages Project. (n.d.). Catalogue of Endangered Languages (ELCat). Retrieved September 15, 2026, from https://www.endangeredlanguages.com/

Hammarström, H., Forkel, R., Haspelmath, M., & Bank, S. (2026). Glottolog 5.3. Max Planck Institute for Evolutionary Anthropology. https://doi.org/10.5281/zenodo.18840935

Hedderich, M. A., Lange, L., Adel, H., Strötgen, J., & Klakow, D. (2021). A survey on recent approaches for natural language processing in low-resource scenarios. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics.

Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 6282–6293. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560

Martinus, L., & Abbott, J. Z. (2019). A focus on neural machine translation for African languages. arXiv. https://arxiv.org/abs/1906.05685

Moran, S., & McCloy, D. (Eds.). (2019). PHOIBLE 2.0. Max Planck Institute for the Science of Human History. https://doi.org/10.5281/zenodo.2626687

Masakhane. (2026). LINGUA Africa: Open call for inclusive AI language projects.

Nekoto, W., Marivate, V., Matsila, T., et al. (2020). Participatory research for low-resourced machine translation: A case study in African languages. Findings of the Association for Computational Linguistics: EMNLP 2020, 2144–2160. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.findings-emnlp.195

Ogueji, K., Zhu, Y., & Lin, J. (2021). Small data? No problem! Exploring the viability of pretrained multilingual language models for low-resourced languages. Proceedings of the 1st Workshop on Multilingual Representation Learning, 116–126. Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.mrl-1.11

Skirgård, H., Haynie, H. J., Blasi, D. E., et al. (2023). Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss. Science Advances, 9(16), eadg6175. https://doi.org/10.1126/sciadv.adg6175

UNESCO. (2002). Atlas of the world’s languages in danger of disappearing (2nd ed.). UNESCO Publishing.

UNESCO. (2010). Atlas of the world’s languages in danger (3rd ed.). UNESCO Publishing.

UNESCO. (n.d.). World Atlas of Languages. Retrieved September 15, 2026, from https://en.wal.unesco.org/

Appendix A

Language categorisation

How languages are classified here, and where the complete list lives.

DimensionValuesSource of the scheme
FamilyNiger-Congo, Afro-Asiatic, Nilo-Saharan, Khoe-Kwadi / Tuu / Kx’a, Austronesian, Indo-European, contact languages, sign languagesGlottolog classification (Section IV.A)
RegionWest, Central, East, Horn, Southern, North AfricaSix-region scheme (Section IV.B)
VitalityVital, Vulnerable, Endangered, Extinct, UnknownUNESCO Atlas and ELCat, kept as separate unreconciled fields
Digital resource tierT1 broad commercial support → T5 no annotated resourceClassification devised for this report
ConfidenceMEASURED, ESTIMATED, STALE, UNKNOWNLabeling scheme in Table 7
Situation groupVital/digitally thin, Shifting, Critical, ReversingCase grouping in Section VII

Table 8 — Categorization scheme applied throughout. See Section IV.A for family definitions, Section IV.B for the regional scheme, and Table 7 for confidence labeling.

Full report
Atilingo Research Report — PDF edition
Download the PDF