January 19, 2026
Language Identification (LID), the task of determining the language of a given text, is a fundamental preprocessing step that shapes the reliability of downstream NLP applications. While recent work has expanded African LID, existing systems remain limited in both language coverage and fine-grained discrimination among closely related languages and varieties. We introduce AfroScope, a unified framework for African LID that includes AfroScope-Data, a dataset covering \(640\) languages, and AfroScope-Models, a suite of strong LID models with broad African language coverage. To address persistent confusions among closely related languages, we propose a hierarchical classification approach that leverages AfroScope-Mirror, a specialized embedding model for targeted disambiguation. This approach improves macro-F1 by \(1.57\) points on the confusable subset compared to our best base model. We further analyze cross-lingual transfer and domain effects, showing how language-family structure, script compatibility, and domain coverage shape LID performance. We position African LID as an enabling technology for large-scale measurement of Africa’s linguistic landscape in digital text and release AfroScope-Dataand AfroScope-Modelsonline.1
Language Identification (LID), the task of determining the language of a given text, is a foundational step in curating multilingual corpora from web crawls [1], [2]. LID errors propagate to downstream stages such as tokenization [3], filtering [4], [5], and data scheduling for multilingual pretraining [6]–[8]. Crucially, LID systems determine not only how reliably each language is predicted but also the scope of identifiable languages. If a language is out of scope, its text is either dropped or misattributed to an in-scope language, distorting corpus composition and downstream evaluation [9], [10].
Although LID is often treated as largely solved [11], [12], recent evaluations show that performance remains uneven across languages and domains [13], [14]. LID for African languages remains especially challenging. Existing benchmarks cover only a small fraction of the continent’s languages, and even within this limited coverage, performance often lags behind that of higher-resource languages. At the same time, major web-crawled corpora for African languages suffer from systematic quality issues [15], including noisy or unusable text, misattributed documents [16], and heavy concentration in religious or translated material that does not adequately reflect everyday language use [17]. These artifacts degrade downstream performance and inflate apparent coverage—a form of representation washing [18] that reinforces disparities in language technology [19]. Recent LID systems [2], [17], including African-focused models [10], have made important progress. However, two central gaps remain: scope, i.e., the set of African languages that can be reliably identified, and granularity, i.e., the ability to distinguish closely related languages and varieties. We address these gaps with AfroScope, a unified framework for African LID. As shown in Figure [fig:main95fig], AfroScope consists of three main contributions:
(i) Coverage-oriented data and models. We introduce AfroScope-Data, a large-scale multilingual dataset curated to expand both language and domain coverage for African LID. AfroScope-Data spans \(640\) languages across multiple orthographies and domains (§3), enabling a more comprehensive assessment of African LID performance, including fine-grained analysis across domains. Using AfroScope-Data, we train AfroScope-Models, a family of LID models that outperform prior African LID baselines across our evaluation setting (§4).
(ii) Hierarchical disambiguation of closely related languages. We show that genetically related and geographically proximate languages are a major source of false positives, especially when labels correspond to closely related varieties or macrolanguage members. To address this, we introduce AfroScope-Mirror, a lightweight contrastive embedding model specialized for frequently confused language groups, and use it in a hierarchical inference procedure that performs targeted disambiguation only when fine-grained separation is needed (§6.1.1). This design improves discrimination among closely related languages while preserving broad-coverage LID.
(iii) Transfer and robustness analysis. Leveraging AfroScope-Data, we analyze LID performance by coverage and domain (§5), and study multilingual transfer effects—including positive transfer and negative interference—as a function of language family structure and script overlap (§6.2). These analyses provide practical guidance for building, evaluating, and curating robust African LID systems.
Figure 1: Language-resource and domain distributions for GlotLID-C, AfroLID, and AfroScope-Data. Each pair of plots shows the language distribution of sentence counts across languages (left) and the corresponding domain composition (right). AfroScope-Data reduces language imbalance and domain concentration relative to the source corpora. Domains accounting for less than \(3\)% of a dataset are grouped under Other.. a — GlotLID-C dataset, b — AfroLID dataset, c — AfroScope-Data
Africa is among the most linguistically diverse regions, spanning many language families and typological profiles [20], [21]. For NLP systems, this diversity manifests in phenomena that directly stress corpus curation and LID, including rich morphology, orthographic variation, and pervasive multilingual practices such as code-switching [13], [22]. In addition, language varieties with fluid boundaries complicate labeling and evaluation [23], [24], particularly for diglossic African languages like Arabic. Recent work has responded with new resources, benchmarks, and African-focused models, which we discuss—along with related challenges—in Appendix 8.1.
Large multilingual corpora frequently contain misattributed text, ambiguous language codes, and other quality issues that disproportionately affect low-resource settings [15]. Prior studies of widely used multilingual resources and pipelines document systematic noise and labeling errors [25]–[27], and emphasize the role of LID quality and preprocessing in mitigating such artifacts [28]. Improving authenticity is therefore central to building reliable and culturally representative language technologies [29]–[31], motivating dataset construction that explicitly controls for coverage, domain diversity, and contamination.
Recent evaluations show that LID coverage for African languages remains narrow and performance lags substantially [14]. Despite progress from FastText-based systems [17], [18], [32] to African-focused transformer-based models [10], [33], [34], as well as methodological advances through contrastive learning [2] and hierarchical approaches [28], targeted efforts for African LID remain necessary to address problems specific to African languages, such as domain variation and closely related varieties [35].
| Dataset | Sent. | Lang. | Family | Script | Domain | ||
|---|---|---|---|---|---|---|---|
| GlotLID-C [17] | \(60{,}682{,}541\) | \(523\) | \(7\) | \(5\) | |||
| AfroLID [10] | \(1{,}682{,}541\) | \(513\) | \(7\) | \(5\) | |||
| SimbaText [36] | \(382{,}541\) | \(101\) | \(5\) | \(4\) | |||
| Train | \(5{,}952{,}575\) | \(640\) | \(8\) | \(7\) | |||
| Dev | \(463{,}875\) | ||||||
| FLORES\(+\) [37] | \(108{,}486\) | \(54\) | \(5\) | \(4\) | |||
| MAFAND [38] | \(54{,}795\) | \(21\) | \(4\) | \(2\) | |||
| SmolSent [39] | \(10{,}872\) | \(53\) | \(6\) | \(4\) | |||
| MCS-350 [28] | \(94{,}894\) | \(141\) | \(6\) | \(3\) | |||
| UDHR [17] | \(6{,}696\) | \(117\) | \(6\) | \(4\) | |||
| CommonLID [14] | \(10{,}696\) | \(25\) | \(3\) | \(3\) | |||
| Test | \(232{,}563\) | \(640\) | \(8\) | \(7\) | |||
| Speech Human Rights Crowdsourcing News Bible Web | |||||||
Building robust LID systems for African languages requires data that is broad across two axes: language coverage and domain coverage. Existing resources cover only a small subset of African languages, leaving many underrepresented or entirely absent and making it difficult to measure true progress. Even among covered languages, sentence counts are highly concentrated in a small number of relatively high-resource languages (Figure 1 (a)). Domain coverage is similarly skewed: religious texts, predominantly Bible translations, account for \(73.10\)% and \(94.70\)% of sentences in GlotLID (Figure 1 (a)) and AfroLID (Figure 1 (b)), respectively. This narrow distribution limits generalization to heterogeneous digital text, including web text, administrative documents, and transcribed speech.
To address these challenges, we curate AfroScope-Data (Figure 1 (c)) using a strategy guided by two objectives: (i) maximizing language coverage to reduce out-of-model cousin errors [15], [40], where text from an unsupported language is misattributed to the closest supported relative; and (ii) increasing domain diversity to mitigate the narrow domain concentration in available African language data.
AfroScope-Data (Table 1) spans \(640\) African languages across eight language families, seven scripts, and eight domains. To the best of our knowledge, AfroScope-Data provides the broadest publicly described coverage for African LID in terms of the joint number of African language labels and domain categories.
We compile AfroScope-Data from publicly described sources, prioritizing datasets that provide sufficient provenance metadata for assigning coarse domain labels. To maximize coverage, we use GlotLID-C [17], AfroLID [10], and SimbaText [36] as training sources, given their breadth and metadata availability
(Table 1).
To mitigate language imbalance, we sample a fixed-size training corpus using temperature smoothing, following prior multilingual practice [9], [18], [41]. For a language \(l\) with corpus fraction \(p_l\), we sample proportional to \(p_l^{\alpha}\), where \(\alpha=0.3\). We apply the same smoothing hierarchically at the domain level: within each language’s budget, a domain \(d\) with intra-language fraction \(p_d\) is sampled proportional to \(p_d^{\alpha}\). This reduces the dominance of high-frequency domains, such as religious texts, and increases the representation of lower-frequency ones.
Through this sampling strategy, AfroScope-Data mitigates imbalance along both axes (Figure 1 (c)). The top-10 languages account for only \(9\)% of sentences, down from \(29\)% in GlotLID-C. Bible-derived content is reduced to \(53.5\)%, compared to \(73.1\)% and \(94.7\)% in GlotLID-C and AfroLID, respectively. We provide more details regarding data curation and sampling in Appendix 9.
AfroScope-Data spans eight high-level groupings: Afro-Asiatic, Austronesian, Creole, Indo-European, Khoe-Kwadi, Mixed language, Niger-Congo, and Nilo-Saharan. This diversity reduces reliance on cues from dominant families such as Niger-Congo and supports evaluation of cross-family generalization. We use this hierarchy in our transfer analyses (§6.2).
| Models | FLORES+ | UDHR (Human Rights) | CommonLID (Web) | SmolSent (Translation) | MAFAND (News) | MCS-350 (Stories) | AfroScope Test | |||||||
| \(n = 54\) | \(n = 117\) | \(n = 25\) | \(n = 53\) | \(n = 21\) | \(n = 141\) | \(n = 640\) | ||||||||
| F1 \(\uparrow\) | FPR \(\downarrow\) | F1 \(\uparrow\) | FPR \(\downarrow\) | F1 \(\uparrow\) | FPR \(\downarrow\) | F1 \(\uparrow\) | FPR \(\downarrow\) | F1 \(\uparrow\) | FPR \(\downarrow\) | F1 \(\uparrow\) | FPR \(\downarrow\) | F1 \(\uparrow\) | FPR \(\downarrow\) | |
| AfroLID | \(61.96\) | \(0.0015\) | \(56.61\) | \(0.0014\) | \(90.51\) | \(0.0010\) | \(58.06\) | \(0.0020\) | \(79.77\) | \(0.0017\) | \(52.53\) | \(0.0011\) | \(69.04\) | \(0.0004\) |
| GlotLID-M | \(95.25\) | \(0.0003\) | \(74.01\) | \(0.0007\) | \(94.33\) | \(0.0003\) | \(80.21\) | \(0.0006\) | 90.87 | \(0.0006\) | \(76.95\) | \(0.0002\) | \(71.85\) | \(0.0003\) |
| ConLID | \(95.35\) | \(0.0003\) | \(74.76\) | \(0.0008\) | \(94.23\) | \(0.0002\) | \(79.58\) | \(0.0006\) | \(90.65\) | \(0.0005\) | 77.56 | \(0.0002\) | \(72.90\) | \(0.0002\) |
| OpenLID | \(83.30\) | \(0.0013\) | \(26.19\) | \(0.0034\) | \(90.94\) | \(0.0009\) | \(47.31\) | \(0.0031\) | \(82.37\) | \(0.0017\) | \(39.83\) | \(0.0020\) | \(20.13\) | \(0.2100\) |
| AfroLID_AS | \(62.11\) | \(0.0024\) | \(59.74\) | \(0.0020\) | \(91.13\) | \(0.0007\) | \(64.75\) | \(0.0024\) | \(79.63\) | \(0.0015\) | \(54.70\) | \(0.0011\) | \(74.38\) | \(0.0003\) |
| Serengeti_AS | \(52.35\) | \(0.0037\) | \(57.72\) | \(0.0023\) | \(92.22\) | \(0.0007\) | \(60.05\) | \(0.0034\) | \(73.28\) | \(0.0027\) | \(45.70\) | \(0.0016\) | \(75.07\) | \(0.0003\) |
| Cheetah_AS | \(58.49\) | \(0.0017\) | \(62.34\) | \(0.0016\) | \(86.89\) | \(0.0005\) | \(59.67\) | \(0.0018\) | \(69.25\) | \(0.0017\) | \(51.53\) | \(0.0008\) | \(79.49\) | \(0.0003\) |
| AfroLID_GL | \(94.61\) | \(0.0003\) | \(72.35\) | \(0.0009\) | \(95.70\) | \(0.0004\) | \(84.08\) | \(0.0007\) | \(83.69\) | \(0.0011\) | \(73.39\) | \(0.0006\) | \(70.09\) | \(0.0004\) |
| Serengeti_GL | \(95.80\) | \(0.0003\) | \(72.82\) | \(0.0009\) | \(96.39\) | \(0.0004\) | \(84.48\) | \(0.0008\) | \(83.98\) | \(0.0012\) | \(75.55\) | \(0.0005\) | \(69.24\) | \(0.0004\) |
| Cheetah_GL | \(92.04\) | \(0.0004\) | \(72.58\) | \(0.0008\) | \(91.81\) | \(0.0004\) | \(78.81\) | \(0.0008\) | \(79.77\) | \(0.0012\) | \(68.65\) | \(0.0006\) | \(75.72\) | \(0.0004\) |
| AfroLID_AF | \(95.06\) | \(0.0004\) | \(74.61\) | \(0.0009\) | \(96.51\) | \(0.0004\) | \(85.07\) | \(0.0007\) | \(87.66\) | \(0.0010\) | \(74.56\) | \(0.0005\) | \(95.82\) | \(0.0001\) |
| + Mirror | \(96.29\) | \(0.0003\) | \(77.53\) | \(0.0007\) | \(95.47\) | \(0.0004\) | \(85.07\) | \(0.0007\) | \(88.69\) | \(0.0009\) | \(74.56\) | \(0.0004\) | \(97.74\) | \(0.0000\) |
| Serengeti_AF | \(95.04\) | \(0.0004\) | \(74.63\) | \(0.0009\) | \(96.51\) | \(0.0004\) | \(85.06\) | \(0.0007\) | \(87.66\) | \(0.0010\) | \(74.56\) | \(0.0005\) | \(95.82\) | \(0.0001\) |
| + Mirror | 97.44 | \(0.0003\) | 77.91 | \(0.0006\) | 96.72 | \(0.0004\) | 85.70 | \(0.0007\) | \(90.13\) | \(0.0008\) | \(76.03\) | \(0.0004\) | 97.87 | \(0.0000\) |
| Cheetah_AF | \(95.62\) | \(0.0003\) | \(76.29\) | \(0.0008\) | \(93.70\) | \(0.0003\) | \(85.59\) | \(0.0007\) | \(85.42\) | \(0.009\) | \(71.77\) | \(0.0005\) | \(96.90\) | \(0.0000\) |
We include seven writing systems: Latin (Latn), Arabic (Arab), Ge’ez (Ethi), N’Ko (Nkoo), Tifinagh (Tfng), Coptic (Copt),
and Vai (Vaii). We explicitly label scripts for each language (language_script), as individual languages may employ multiple writing systems (e.g., gof, ttq), enabling script-aware evaluation and analysis of
orthographic variation.
To analyze domain effects, we first adopt the source-domain categorization from GlotLID [17], which groups sources into coarse categories such as bible, news, web, government, and social text. We extend this mapping to sources absent from the original GlotLID taxonomy, using the origin metadata each dataset provides. This yields a compact set of domains: Speech, Human Rights, Crowdsourcing, News, Bible, Web, Government, and Other. We assign each sentence a domain label through this source-level mapping and use the labels for controlled evaluation by domain; we discuss domain effects in detail in §5.
We evaluate on six external Evaluation sets spanning diverse domains and language coverage: FLORES+ [37],
UDHR [17], MAFAND [38], SmolSent [39], MCS-350 [28], and CommonLID [14]. We also evaluate on a held-out AfroScope-Data test split. Together, these benchmarks allow us to assess broad-coverage African LID and robustness across domains.
Because several African LID resources are derived from overlapping public sources, we explicitly measure residual overlap between Train and Evaluation data. We use a strict 4-gram contamination criterion: a test sentence is marked as contaminated if all of its 4-grams appear within a single training sentence. Table 3 reports contamination rates across three training corpora and the six external evaluation sets. AfroScope-Data exhibits no detected contamination on any of the Evaluation sets under this criterion, whereas the AfroSimba and GlotLID-C corpora show non-trivial overlap, particularly on MCS-350 (\(14.15\)% with GlotLID-C) and MAFAND (\(8.13\)% with GlotLID-C). We return to the relationship between contamination and downstream model performance in §5.
| Eval Set | AfroSimba | Glotlid-C | AfroScope-DATA | |||
|---|---|---|---|---|---|---|
| 2-3(lr)4-5(lr)6-7 | Cont. | Lang. | Cont. | Lang. | Cont. | Lang. |
| FLORES+ | \(0.00\%\) | \(0\) | \(0.00\%\) | \(0\) | \(0.00\%\) | \(0\) |
| UDHR | \(\mathbf{1.75\%}\) | \(33\) | \(0.97\%\) | \(10\) | \(0.00\%\) | \(0\) |
| CommonLID | \(1.41\%\) | \(16\) | \(\mathbf{7.19\%}\) | \(18\) | \(0.00\%\) | \(0\) |
| SMOL | \(0.00\%\) | \(1\) | \(\mathbf{0.04\%}\) | \(16\) | \(0.00\%\) | \(0\) |
| MAFAND | \(0.37\%\) | \(12\) | \(\mathbf{8.13\%}\) | \(18\) | \(0.00\%\) | \(0\) |
| MCS-350 | \(1.47\%\) | \(84\) | \(\mathbf{14.15\%}\) | \(97\) | \(0.00\%\) | \(0\) |
Following [17], we evaluate each model on the full benchmark rather than filtering examples to the model’s supported language set. This setting reflects more realistic LID deployment conditions, where systems encounter text in languages outside their training label space and must avoid incorrectly assigning such text to an in-scope language. It also makes coverage differences visible: models with narrower label spaces are penalized when they force unsupported languages into related or superficially similar supported labels.
We report macro-F1, which averages performance across languages, together with False Positive Rate (FPR), defined as: \(\text{FPR} = \frac{\text{FP}}{\text{FP} + \text{TN}}\), where FP is the number of false positives and TN is the number of true negatives. Macro-F1 measures average per-language identification quality, while FPR captures how often a model falsely attributes text to a language it does not belong to. This distinction is important for African LID, where unsupported or closely related languages can otherwise be misclassified as higher-resource relatives.
We compare against a diverse set of LID systems, ranging from FastText-based classifiers to transformer-based African language models.
We evaluate three existing FastText-based LID systems: GlotLID-M [17], OpenLID [18] and ConLID [2]. ConLID incorporates supervised contrastive learning to improve robustness on out-of-domain data.
We evaluate three transformer-based models developed for African languages: AfroLID [10], Serengeti [33], and Cheetah [34]. To isolate the effect of training data, we fine-tune these models on each Train source where applicable, and refer to the resulting models collectively as AfroScope-Models. Because SimbaText is substantially smaller than the other training sources, we merge it with AfroLID to form a combined training source, AfroSimba. Fine-tuning hyperparameters are provided in Appendix 10.
Table 2 reports macro-F1 and FPR across all external Evaluation sets for the full set of baselines and AfroScope-Modelsvariants. AfroScope-Modelsachieve the strongest overall performance, leading on five of seven benchmarks, with the largest gains appearing on benchmarks with broader language coverage. The strongest improvements occur on the two broadest evaluation sets, AfroScope-DataTest (\(640\) languages) and UDHR (\(117\) languages), where the best SerengetiAF\(+\) Mirror variant obtains \(97.87\) and \(77.91\) macro-F\(1\), respectively, with FPRs of \(0.0000\) and \(0.0006\). Datasets with narrower language coverage, such as FLORES+, CommonLID, and SmolSent, show the same general pattern but with smaller margins. GlotLID-M and ConLID, both trained on the GlotLID-C corpus, achieve the strongest results only on MAFAND (\(90.87\) and \(90.65\) macro-F1) and MCS-350 (\(77.56\) macro-F1). These are also the evaluation datasets for which the GlotLID-C corpus exhibits the highest contamination, suggesting that evaluation overlap may contribute to their advantage on these benchmarks rather than reflecting broader generalization.
Coverage is a major driver of performance differences across evaluation sets. Languages absent from a model’s training label space are vulnerable to forced misattributions, where the model assigns them to a related or superficially similar supported language, thereby inflating FPR. Across the seven evaluation sets, the GlotLID-C corpus label space lacks coverage for \(171\) evaluation-language entries while AfroScope-Datacovers all but \(31\). On datasets where both corpora cover nearly all evaluation languages, such as FLORES+, MAFAND, and SmolSent, model performance clusters within a relatively narrow macro-F1 range. In contrast, on broader evaluation sets with larger coverage gaps, such as UDHR and MCS-350, AfroScope-Modelsvariants achieve an average macro-F1 of \(75.25\) compared to \(67.25\) for non-AfroScope-Databaselines, while average FPR decreases from \(0.00101\) to \(0.00067\).
| Models | AfroScope Test | |
| \(F1\uparrow\) | \(FPR\downarrow\) | |
| GlotLID-M | \(69.08\) | \(0.0002\) |
| SerengetiAF | \(97.71\) | \(0.0001\) |
| Gemini-Pro | \(36.40\) | \(0.0069\) |
| GPT-5.4-mini | \(45.30\) | \(0.0043\) |
| Claude-4.7-OPUS | \(47.40\) | \(0.0041\) |
Figure 2 shows that model performance is strongly shaped by the amount and distribution of domain-specific training data. High-resource domains such as Bible, News, Crowdsourcing, and Government generally achieve strong per-language macro-F1 scores, with most languages clustered near the top of the distribution. In contrast, lower-resource or more heterogeneous domains exhibit weaker and less stable behavior. Human Rights contains only a few training examples across \(60\) languages, and its per-language results include a severe lower tail, with several languages receiving near-zero macro-F1. This pattern helps explain the relatively low performance on UDHR, a human-rights benchmark. Similarly, the Web domain has moderate training coverage but noticeably weaker lower-tail performance, suggesting that domain diversity and text heterogeneity also affect robustness beyond raw example count alone. Overall, these results show that LID performance depends on both language coverage and domain coverage: domains that are well represented in training generalize more reliably, while sparse or unevenly distributed domains produce brittle performance for some languages.
Given that LLMs now set the state of the art across many NLP tasks, we evaluate three frontier LLMs on AfroScope-Datatest: Gemini-Pro, Claude-4.7-OPUS, and GPT-5.4-mini. Prompt templates, candidate-label formatting, and other evaluation details are provided in Appendix 11. As shown in Table 4, all three LLMs lag substantially behind dedicated LID systems: the strongest, Claude-4.7-OPUS, reaches only \(47.40\) F1 (FPR \(0.0041\)), compared to \(69.08\) for GlotLID-M and \(97.71\) for AfroScope-Models. This gap suggests that despite their broad multilingual capabilities, general-purpose LLMs are not a substitute for purpose-built LID models on African languages.
In this section, we analyze the patterns behind the aggregate results in Table 2. We focus primarily on AfroScope-Data test because it has the broadest language coverage, while using external evaluation sets to test whether the same patterns recur across domains and benchmarks. Our analysis addresses two questions: (i) how does confusability among closely related languages affect performance? and (ii) which languages synergize or interfere most with one another, and how do language family relationships and script compatibility shape these effects?
We find that label confusability is a primary failure mode: the repeated false-positive assignment between a focus language and a small set of related or geographically proximate labels. Table 5 previews representative confusion patterns identified across the evaluation sets. Such cases are especially common across ISO macrolanguage groupings and closely related languages or varieties that share substantial lexical and orthographic overlap. We find that these confusions recur in at least four of six evaluation sets, indicating the errors are stable across datasets rather than isolated benchmark artifacts. Among macrolanguage varieties, Arabic (ara: aeb, ary, arz) and Fulah (ful: fub, fuf, fuv) are representative. High-confusion pairs are often genetically related and geographically proximate, such as Lamnso (lns) and Tikar (tik) in Cameroon, and Anii (blo) and Lukpa (dop) in Togo. These groupings account for a disproportionate share of false positives, suggesting that models struggle to discriminate fine-grained varieties. Together, these patterns indicate that many errors reflect genuine linguistic similarity and label granularity rather than random noise. We provide further examples in Appendix 12.1.
| Lang. | FP. | FC. | TC. | |
|---|---|---|---|---|
| ara | \(0.093\) | \(2281\) | aeb, apd, arq, ary | |
| din | \(0.040\) | \(497\) | dib, dik, dip, dks | |
| ful | \(0.010\) | \(414\) | ffm, fub, fuc, fue | |
| lns | \(0.281\) | \(2577\) | tik, pfe, agq | |
| blo | \(0.199\) | \(1748\) | kbp, dop, bib | |
| kin | \(0.153\) | \(4323\) | run, swh, nnb |
To address these systematic confusions, we introduce AfroScope-Mirror, a lightweight contrastive embedding model used inside a hierarchical inference procedure. The base LID model first produces a broad-coverage prediction. When the prediction falls within a predefined confusable group and the model’s confidence is below a threshold \(\tau\), AfroScope-Mirrorperforms a group-specific disambiguation step using specialized embeddings. We build AfroScope-Mirroron top of AfroLIDAF and SerengetiAF, our strongest base models on the external evaluation sets, and train it with Mirror-BERT [42], an unsupervised contrastive learning objective that pulls semantically similar representations together while pushing unrelated ones apart. Detailed training procedures and hyperparameters are in Appendix 11.
Figure 3 compares the embedding spaces produced by SerengetiAF and SerengetiAF \(+\) Mirror, visually illustrating clearer separation among confusable labels and tighter within-label clustering after specialization. We evaluate this strategy on the \(30\) confusable languages in Table 12, including focus languages and their top confusion partners, across confidence thresholds. On the full set, the average macro-F1 gain is +\(1.57\) points. As shown in Table 2, adding AfroScope-Mirror yields consistent gains across the external benchmarks: SerengetiAF \(+\) Mirror improves over its base model on all seven evaluation sets (e.g., +\(2.40\) on FLORES+, +\(3.28\) on UDHR, +\(2.47\) on Mafand), and AfroLIDAF \(+\) Mirror improves on five of seven datasets. At the language level (Table 6), we observe improvements for closely related varieties within macro-language groupings, such as Arabic (ara) varieties aeb, arq, and ary, which improve by +\(0.63\), +\(0.79\), and +\(0.31\), respectively, as well as Swahili (swa) variety swc (+\(0.55\)). We also observe improvements for confusable regional pairs, including Kinyarwanda (kin) and run (+\(0.71\)). However, a small number of languages decline, including niq (\(-0.10\)) and kbp (\(-0.14\)), suggesting that hierarchical routing can introduce unnecessary complexity when the base model is already reliable for a given label.
| Group | Lang | Baseline | \(\tau=0.75\) | \(\tau=0.85\) | \(\tau=0.95\) | |||
|---|---|---|---|---|---|---|---|---|
| 4-5 (lr)6-7 (lr)8-9 | \(\boldsymbol{\Delta}\) | F1 | \(\boldsymbol{\Delta}\) | F1 | \(\boldsymbol{\Delta}\) | F1 | ||
| ara | aeb | \(94.50\) | \(+0.18\) | \(94.68\) | \(\mathbf{+0.63}\) | \(\mathbf{95.13}\) | \(+0.50\) | \(95.00\) |
| arq | \(98.89\) | \(+0.32\) | \(99.21\) | \(+0.64\) | \(99.53\) | \(\mathbf{+0.79}\) | \(\mathbf{99.68}\) | |
| ary | \(96.05\) | \(+0.11\) | \(96.16\) | \(\mathbf{+0.31}\) | \(\mathbf{96.35}\) | \(+0.13\) | \(96.17\) | |
| swa | swc | \(94.21\) | \(+0.34\) | \(94.55\) | \(+0.07\) | \(94.28\) | \(\mathbf{+0.55}\) | \(\mathbf{94.77}\) |
| kin | run | \(95.34\) | \(+0.03\) | \(95.37\) | \(+0.03\) | \(95.37\) | \(\mathbf{+0.71}\) | \(\mathbf{96.05}\) |
| Avg | – | \(+0.19\) | – | \(+0.33\) | – | \(\mathbf{+0.54}\) | – | |
Given limited and skewed resources, effective African LID depends on cross-lingual transfer. We investigate which languages synergize or interfere most, and how language family and script compatibility shape these effects. To quantify this, we adapt the Bilingual Transfer Score (BTS) [43] to language identification. We compute BTS using macro-F1, and cross-entropy loss across \(30\) representative African languages spanning five family clusters.
Figure 7 shows that cross-lingual transfer in African LID is governed by shared script and fine-grained language relatedness rather than coarse family membership. Co-training yields the largest F1 gains for low-resource targets with weak monolingual baselines (e.g., fuv, dyu, bam; F1 BTS up to +\(0.22\)), consistent with baseline rescue rather than family-specific transfer. Critically, the closest relationships are not uniformly beneficial: mutually intelligible varieties such as the Arabic cluster (arz, ary, aeb) and Mande pair bam–dyu improve F1 but degrade probability calibration, with loss-based BTS as low as -\(0.24\). A regression across all pairs confirms that script and this confusable-cousin relationship—not language family—are the dominant predictors of transfer; family membership alone is not significant. This accuracy–calibration trade-off motivates hierarchical disambiguation. We report full per-family analysis in Appendix 12.2.
We introduce AfroScope, a unified framework for African LID that combines broad-coverage data, strong LID models, targeted disambiguation, and linguistic analysis. We present AfroScope-Data, a large-scale dataset spanning \(640\) language labels, which we use to train AfroScope-Models, a family of African LID models that outperform prior African-focused baselines across internal and external evaluations.
To address persistent confusions among closely related languages and varieties, we propose a hierarchical inference approach based on AfroScope-Mirror, a specialized embedding model for frequently confused language groups. On the full set of identified confusion groups, this approach improves macro-F1 by \(+1.57\) points on average. Finally, our transfer analysis shows that African LID transfer is shaped by script, fine-grained family relatedness and confusability than by coarse family membership. These relationships can improve low-resource performance while degrading calibration, motivating targeted disambiguation with AfroScope-Mirror. We hope AfroScope, AfroScope-Data, and AfroScope-Modelssupport reliable African NLP and future work on fine-grained varieties, domain shifts, and mixed-language text.
We note several limitations of the current study.
Mixed-language and code-switched text. Our formulation treats each instance as belonging to a single language label. This does not fully capture important phenomena in Africa’s linguistic landscape, including code-switching, mixed-language documents, translanguaging practices, and contact varieties such as pidgins and creoles. Extending AfroScopeto multi-label, document-level, or span-level language identification is an important direction for future work. Relatedly, the granularity at which languages are annotated is uneven across our data sources, and in some cases closely related varieties or distinct languages may be conflated under a single label. A more careful audit of these labeling decisions, ensuring that language distinctions are drawn consistently and appropriately, is needed but lies beyond the scope of this work.
Language metadata and classification choices. AfroScope-Datarelies on external catalogs, primarily Ethnologue, to assign language identifiers, genealogical groupings, and script metadata. Alternative resources, such as Glottolog, may differ in classification, naming, macrolanguage treatment, and family hierarchy. These choices may affect analyses that depend on genealogical proximity or script compatibility. Future work should evaluate sensitivity to alternative metadata sources and provide mappings across catalog standards.
Confidence-based routing and calibration. Our hierarchical disambiguation method relies on model confidence to decide when to invoke the group-specific refinement step. While this improves performance on many confusable labels, gains are not uniform across languages or thresholds, and some cases exhibit degradation. Improving probability calibration and learning a routing policy, rather than relying on fixed confidence thresholds, may further increase robustness.
Domain and source imbalance. Although AfroScope-Dataimproves domain diversity relative to prior African LID resources, some domains remain substantially smaller or more heterogeneous than others. As a result, performance may still be brittle for domains with limited training data or high internal variation, such as web text, human-rights documents, or informal user-generated content. Future work should expand naturally occurring data in underrepresented domains and evaluate robustness under domain shift.
The following appendices provide comprehensive supplementary material supporting the main findings of this work. We include an expanded literature review, detailed descriptions of the datasets and models, a more in-depth discussion of the experimental setup, and additional analyses and discussions.
§8: More in depth literature review
§9: Data Curation & Preprocessing
§10: Baseline Models
§12: Discussions
Many African languages are agglutinative (e.g., Swahili, Zulu), complicating tokenization and parsing [22], [23], while everyday communication often involves code-switching with English and other languages [44].
Datasets like AfroCS-xs demonstrate that small, high-quality resources can boost LLM performance on code-switched tasks [45]. Models such as Afrolid, SERENGETI and Cheetah show the feasibility of massively multilingual modeling for African languages [10], [33], [34]. Benchmarks including Sahara [46], SimbaBench [36], IrokoBench [47], and AfroBench [29].
Studies of multilingual corpora such as ParaCrawl [25], WikiMatrix [26], and mC4 [27] reveal widespread errors, mislabeled data, and ambiguous language codes [15]. For instance, [16] showed that many FastText embeddings were heavily mislabeled with words from other languages. Poor data authenticity directly undermines cultural representation, as inaccurate or mislabeled data fails to capture the nuanced linguistic and cultural contexts inherent in low-resource languages [30]. Addressing these quality issues is essential for building reliable language technologies [29], [31].
We deduplicate sentences across all sources to ensure that each example appears at most once in AfroScope-Data. We then partition the data into train, dev, and test splits, performing deduplication prior to splitting so that no sentence is shared across splits. We ensure that each sentence is written in the correct script, based on the writing system databases of Ethnologue [48]. When temperature sampling we set the maximum sentences to \(10\)M sentences. The resulting splits contain \(5{,}952{,}573\) sentences for train, \(463{,}875\) for dev, and \(232{,}563\) for test. The full list of languages supported by AfroScope, together with the number of sentences per language, is provided in Tables 15–18.
Below, we describe the sources used to construct AfroScope-Data.
We extract 523 African languages from GlotLID-C, a collection spanning 2,099 languages globally [17].
A manually curated multi-domain web dataset covering 513 African languages [10].
Speech-derived text data spanning 103 African languages, originally collected for speech and language identification [36].
Professionally translated sentences across \(200\) languages, providing a clean and standardized benchmark, particularly for low-resource languages [37].
Manually audited data from news, Wikipedia, and religious texts across 55 African languages [38].
Multilingual benchmark dataset consisting of professionally translated sentence- and document-level data for 115 LRLs [39].
Parallel children’s stories across 151 African languages, drawn from a multilingual collection of 50k texts in over 350 languages [28].
The Universal Declaration of Human Rights, translated into a wide range of languages [17].
Professionally translated parallel data for 115 under-represented languages [14].
We assign domains by matching keywords found in the metadata associated with each sentence, following the categorization scheme from [17]. Table 8 lists the specific keyword mappings.
| Sample Metadata |
|---|
| Bible-aar_line94 |
| CommonVoice v11 |
| Masakhanews |
| GlotStoryBook |
| CC100_zu.txt.tsv_f17_line40162 |
| JW-zul_line3295 |
| gov-za |
| Open Subtitles |
| KDE4 |
| Vuk’uzenzele |
| Domain | Associated Keywords |
|---|---|
| Speech | Speech, CommonVoice, TTS, Audio |
| Human Rights | Human Rights |
| Government | Human Rights, Autshumato, Legal, GOV, Parliament, Gazette |
| CrowdSource | Flores, NLB, mt560, Tatoeba, UD, ai4d, lti, Benchmark, Human, Madar, iadd |
| News | News, xlsum, Vukuzenzele, CBC, BBC, Afriqa, Masakha, Goud |
| Bible | Bible, JW, Scripture, Religion |
| Web | Oscar, CC, CommonCrawl, Web, Dialect, Social, Forum |
| Other | Health, Covid, Medical, Med, Tanzil, PBC, Quran, Story, Stories, Fiction, Bloom, Lyrics |
3pt
To isolate the effect of training data, we train or fine-tune comparable model architectures on each Train source where applicable. This allows us to evaluate whether the gains come from the model architecture, the training corpus, or their interaction. Because SimbaText is substantially smaller than the other training sources, we merge it with AfroLID to form a combined training source, AfroSimba. AfroLID and Serengeti are XLM-RoBERTa variants, and Cheetah is a T5-based model.
We use the hyperparameters in Table 9 for AfroLID [10] and Serengeti [33] (both XLM-R variants), and Table 10 for Cheetah [34].
Figure 4 shows the prompt template used to elicit language predictions from the LLMs. We provide the candidate ISO 639-3 labels in the prompt and instruct the model to return only the language code. To reduce inference cost we downsample each language examples to 100 examples per language.
| argument | description | value |
|---|---|---|
| -max_seq_length | max input sequence length | 128 |
| -per_device_train_batch_size | training batch size (per device) | 64 |
| -learning_rate | learning rate | 2e-5 |
| -num_train_epochs | number of training epochs | 10 |
| -metric_for_best_model | evaluation metric | f1 |
| argument | description | value |
|---|---|---|
| -max_target_length | max target sequence length | 128 |
| -per_device_train_batch_size | training batch size (per device) | 32 |
| -learning_rate | learning rate | 5e-5 |
| -num_train_epochs | number of training epochs | 10 |
| -metric_for_best_model | evaluation metric | f1 |
To investigate languages with high FPR, we isolate languages with scores below F1 90 across all evaluation datasets—and identify the top three most frequent misclassifications for each to form confusion groups. This analysis yields \(13\) distinct groups comprising \(30\) languages in total. We find that these confusions primarily stem from either macrolanguage structures (e.g., ful vs. fub) or geographic proximity (e.g., bsq vs. bas). Table 12 details the composition of all confusion groups, and we report the corresponding performance improvements for each individual language.
We follow the procedures from the official github repository for Mirror-BERT2. We use SerengetiAF for this experiment as it being our best performing model. Table 11 shows the hyperparamerters we use to train AfroScope-Mirror. We see in Figure 5 that the label spaces of SerengetiAF + Mirror shows much better separation in the embedding spaces than SerengetiAF.
| argument | description | value |
|---|---|---|
| -epoch | number of training epochs | 1 |
| -train_batch_size | training batch size | 200 |
| -learning_rate | learning rate | 2e-5 |
| -max_length | max sequence length | 50 |
| -infoNCE_tau | InfoNCE temperature (\(\tau\)) | 0.04 |
| -dropout_rate | dropout rate | 0.0 |
| -drophead_rate | drophead rate | 0.05 |
| -random_span_mask | length of random span mask | 5 |
| -agg_mode | aggregation mode | cls |
We adapt the Bilingual Transfer Score (BTS) [43] to language identification. For a source \(s\) and target \(t\), we compare a monolingual model trained on target data \(\mathcal{D}_t\) against a bilingual model trained on \(\mathcal{D}_t \cup \mathcal{D}s\), evaluated on a held-out target test set. For cross-entropy loss, where lower is better, we define: \[\mathrm{BTS}^{\mathcal{L}}_{s \rightarrow t} = \frac{ \mathcal{L}_t(\mathcal{D}_t) - \mathcal{L}_t(\mathcal{D}_t \cup \mathcal{D}_s) }{ \mathcal{L}_t(\mathcal{D}_t) + \epsilon }. \label{eq:bts95loss}\tag{1}\] so that positive values indicate improvement. For macro-F1, where higher is better, we define: \[\mathrm{BTS}^{F1}_{s \rightarrow t} = \frac{ F1_t(\mathcal{D}_t \cup \mathcal{D}_s) - F1_t(\mathcal{D}_t) }{ F1_t(\mathcal{D}_t) + \epsilon }. \label{eq:bts95f1}\tag{2}\]
We select \(30\) languages spanning five clusters. To populate a range of genealogical distances, we group eligible languages (those with at least \(10{,}000\) sentences) by their shared family hierarchy and retain clusters that contain documented confusable pairs, drawn from our false-positive analysis (Appendix 12.1). This yields clusters of mutually intelligible varieties—the Arabic dialects (arz, ary, aeb), Fula varieties (fub, fuv, fuh, fuq, fue), and the Mande pair (bam, dyu)—alongside more distant same-family languages and unrelated outgroups (plt, afr, sag). Figure 6 shows the fine-grained relatedness of the selected languages across the family hierarchy.
Each language contributes \(1{,}000\) target sentences and a held-out test set; bilingual models receive \(200\) background sentences from every other language so all models share one label space. We fine-tune SerengetiAF for the \(30\) monolingual baselines (median over three seeds) and the \(435\) language pairs (one seed), yielding a \(30\) X \(30\) directional transfer matrix per metric.
The largest F1-BTS values occur for the weakest baselines—fuv (\(0.62\)), fuq (\(0.53\)), dyu (\(0.69\)), bam (\(0.73\))—which improve from nearly any source (bam\(\rightarrow\)dyu \(=+0.22\), fub\(\rightarrow\)fuv \(=+0.16\) ; Table 13). F1-BTS correlates with target headroom (\(r=0.60\)), indicating low-resource rescue rather than family-specific transfer; strong baselines (kin, xho, tso, tsn, all \(>0.98\)) show near-zero BTS.
The clearest signal is in loss-BTS. The Arabic varieties form a block of negative loss-BTS (aeb\(\rightarrow\)ary \(=-0.24\), aeb\(\rightarrow\)arz \(=-0.18\), ary\(\rightarrow\)aeb \(=-0.15\)), and the Mande pair replicates this (dyu\(\rightarrow\)bam \(=-0.17\)). All show positive F1-BTS: co-training mutually intelligible varieties picks the right label but spreads probability mass across the cousin, raising cross-entropy. This trade-off is visible in Figure 7 (a), where cousin pairs sit in the mutual-gain region under F1 but the mutual-loss region under cross-entropy. Bambara (bam) and Jula (dyu) are adjacent Manding varieties spoken across Mali, Côte d’Ivoire, and Burkina Faso [48], so this interference is expected.
Substantial transfer is confined to same-script pairs: \(78\%\) of F1 changes with \(|\mathrm{BTS}|\geq0.05\) and all calibration changes with \(|\mathrm{BTS}|\geq0.10\) occur within a script (Table 14). Controlling for headroom and cousin status, shared script independently predicts larger F1-transfer magnitude (\(\beta=+0.005\)). The exception is low-resource rescue, where weak Fula targets improve from any source. Distinctive scripts with strong baselines, like the Ethiopic pair amh–tir (\(>0.99\)), show negligible transfer (mean \(|\mathrm{BTS}^{\mathcal{L}}|=0.006\)), as orthographic distinctiveness leaves no confusion to resolve.
Different-family pairs show loss-BTS near zero (distance-\(5\) mean \(=+0.002\), \(n=472\)). Regressing loss-BTS on genealogical distance \(d\) and \(d^2\) yields a significant inverted-U (\(\beta_d=+0.015\), \(\beta_{d^2}=-0.002\)): distance-\(0\) cousins are uniquely negative (mean \(=-0.029\)), all greater distances near zero. F1-BTS shows no distance pattern, masked by ceiling saturation.
| Group | Lang | Baseline | \(\tau=0.75\) | \(\tau=0.85\) | \(\tau=0.95\) | |||
|---|---|---|---|---|---|---|---|---|
| 4-5 (lr)6-7 (lr)8-9 | F1 | \(\boldsymbol{\Delta}\) | F1 | \(\boldsymbol{\Delta}\) | F1 | \(\boldsymbol{\Delta}\) | ||
| aar | aar | \(97.10\) | \(\mathbf{98.24}\) | \(\mathbf{+1.14}\) | \(\mathbf{98.24}\) | \(\mathbf{+1.14}\) | \(\mathbf{98.24}\) | \(\mathbf{+1.14}\) |
| bsq | bas | \(94.23\) | \(\mathbf{95.36}\) | \(\mathbf{+1.13}\) | \(\mathbf{95.40}\) | \(\mathbf{+1.17}\) | \(\mathbf{95.46}\) | \(\mathbf{+1.23}\) |
| ara | aeb | \(94.50\) | \(94.68\) | \(+0.18\) | \(\mathbf{95.13}\) | \(\mathbf{+0.63}\) | \(95.00\) | \(+0.50\) |
| ara | \(92.55\) | \(92.84\) | \(+0.29\) | \(92.99\) | \(+0.44\) | \(\mathbf{93.12}\) | \(\mathbf{+0.57}\) | |
| arq | \(98.89\) | \(99.21\) | \(+0.32\) | \(99.53\) | \(+0.64\) | \(\mathbf{99.68}\) | \(\mathbf{+0.79}\) | |
| ary | \(96.05\) | \(96.16\) | \(+0.11\) | \(\mathbf{96.35}\) | \(\mathbf{+0.31}\) | \(96.17\) | \(+0.13\) | |
| arz | \(93.46\) | \(95.70\) | \(+2.24\) | \(\mathbf{96.23}\) | \(\mathbf{+2.77}\) | \(98.91\) | \(+5.45\) | |
| kin | kin | \(92.90\) | \(93.17\) | \(+0.27\) | \(93.17\) | \(+0.27\) | \(\mathbf{93.57}\) | \(\mathbf{+0.66}\) |
| run | \(95.34\) | \(95.37\) | \(+0.03\) | \(95.37\) | \(+0.03\) | \(\mathbf{96.05}\) | \(\mathbf{+0.71}\) | |
| swh | \(99.80\) | \(\mathbf{99.90}\) | \(\mathbf{+0.10}\) | \(\mathbf{99.90}\) | \(\mathbf{+0.10}\) | \(\mathbf{99.90}\) | \(\mathbf{+0.10}\) | |
| nnb | \(98.58\) | \(\mathbf{98.72}\) | \(\mathbf{+0.14}\) | \(\mathbf{98.72}\) | \(\mathbf{+0.14}\) | \(\mathbf{98.72}\) | \(\mathbf{+0.14}\) | |
| lns | tik | \(99.60\) | \(\mathbf{99.70}\) | \(\mathbf{+0.10}\) | \(\mathbf{99.70}\) | \(\mathbf{+0.10}\) | \(\mathbf{99.70}\) | \(\mathbf{+0.10}\) |
| pfe | \(96.83\) | \(\mathbf{96.96}\) | \(\mathbf{+0.13}\) | \(\mathbf{96.96}\) | \(\mathbf{+0.13}\) | \(\mathbf{96.96}\) | \(\mathbf{+0.13}\) | |
| agq | \(95.16\) | \(\mathbf{95.29}\) | \(\mathbf{+0.14}\) | \(\mathbf{95.29}\) | \(\mathbf{+0.14}\) | \(\mathbf{95.29}\) | \(\mathbf{+0.14}\) | |
| blo | kbp | \(98.51\) | \(\mathbf{98.64}\) | \(\mathbf{+0.13}\) | \(\mathbf{98.64}\) | \(\mathbf{+0.13}\) | \(\mathbf{98.64}\) | \(\mathbf{+0.13}\) |
| dop | \(99.17\) | \(\mathbf{99.31}\) | \(\mathbf{+0.14}\) | \(\mathbf{99.31}\) | \(\mathbf{+0.14}\) | \(\mathbf{99.31}\) | \(\mathbf{+0.14}\) | |
| bib | \(97.41\) | \(\mathbf{97.54}\) | \(\mathbf{+0.13}\) | \(\mathbf{97.54}\) | \(\mathbf{+0.13}\) | \(\mathbf{99.35}\) | \(\mathbf{+1.94}\) | |
| maf | maf | \(95.37\) | \(\mathbf{99.50}\) | \(\mathbf{+4.13}\) | \(\mathbf{99.50}\) | \(\mathbf{+4.13}\) | \(\mathbf{99.50}\) | \(\mathbf{+4.13}\) |
| mcn | \(96.05\) | \(98.17\) | \(+2.12\) | \(98.37\) | \(+2.32\) | \(\mathbf{99.18}\) | \(\mathbf{+3.13}\) | |
| meq | \(95.26\) | \(\mathbf{99.86}\) | \(\mathbf{+4.60}\) | \(\mathbf{99.86}\) | \(\mathbf{+4.60}\) | \(\mathbf{99.86}\) | \(\mathbf{+4.60}\) | |
| mfz | \(95.74\) | \(97.79\) | \(+2.05\) | \(97.89\) | \(+2.15\) | \(\mathbf{98.84}\) | \(\mathbf{+3.10}\) | |
| mpe | \(96.19\) | \(98.56\) | \(+2.37\) | \(98.76\) | \(+2.57\) | \(\mathbf{98.96}\) | \(\mathbf{+2.77}\) | |
| mug | \(96.60\) | \(98.90\) | \(+2.30\) | \(99.30\) | \(+2.70\) | \(\mathbf{99.70}\) | \(\mathbf{+3.10}\) | |
| myx | \(94.96\) | \(\mathbf{95.86}\) | \(\mathbf{+0.90}\) | \(\mathbf{95.86}\) | \(\mathbf{+0.90}\) | \(\mathbf{95.86}\) | \(\mathbf{+0.90}\) | |
| din | din | \(70.51\) | \(\mathbf{73.75}\) | \(\mathbf{+3.24}\) | \(72.96\) | \(+2.44\) | \(\mathbf{73.75}\) | \(\mathbf{+3.24}\) |
| ful | fub | \(92.28\) | \(95.62\) | \(+3.34\) | \(96.33\) | \(+4.05\) | \(\mathbf{96.59}\) | \(\mathbf{+4.31}\) |
| sot | sot | \(92.83\) | \(\mathbf{94.08}\) | \(\mathbf{+1.25}\) | \(\mathbf{94.08}\) | \(\mathbf{+1.25}\) | \(\mathbf{94.08}\) | \(\mathbf{+0.25}\) |
| ibo | ibo | \(98.58\) | \(\mathbf{99.53}\) | \(\mathbf{+0.95}\) | \(\mathbf{99.53}\) | \(\mathbf{+0.95}\) | \(\mathbf{99.53}\) | \(\mathbf{+0.95}\) |
| wol | wol | \(88.13\) | \(\mathbf{90.35}\) | \(\mathbf{+2.22}\) | \(\mathbf{90.35}\) | \(\mathbf{+2.22}\) | \(\mathbf{90.35}\) | \(\mathbf{+2.22}\) |
| swa | swc | \(94.21\) | \(94.55\) | \(+0.34\) | \(94.28\) | \(+0.07\) | \(\mathbf{94.77}\) | \(\mathbf{+0.55}\) |
| Average | – | – | \(\mathbf{+1.22}\) | – | \(\mathbf{+1.29}\) | – | \(\mathbf{+1.57}\) | |
| Source | Target | BTS\(^{F1}\) | BTS\(^{\mathcal{L}}\) | Cousin |
|---|---|---|---|---|
| bam | dyu | +0.22 | \(-\)0.08 | ✔ |
| fub | fuv | +0.16 | \(-\)0.11 | ✔ |
| nnb | fuv | +0.14 | \(-\)0.09 | |
| dyu | bam | +0.14 | \(-\)0.17 | ✔ |
| tsn | fuv | +0.13 | \(-\)0.08 | |
| xho | fuv | +0.13 | \(-\)0.08 | |
| fuh | fuq | +0.12 | \(-\)0.02 | |
| afr | fuv | +0.12 | +0.01 | |
| fub | fuq | +0.12 | +0.03 | |
| tso | fuv | +0.12 | \(-\)0.12 |
| Script pair | \(n\) | \(\overline{|\mathrm{BTS}^{F1}|}\) | \(\overline{|\mathrm{BTS}^{\mathcal{L}}|}\) |
|---|---|---|---|
| Latn+Latn | 506 | 0.015 | 0.012 |
| Arab+Latn | 184 | 0.010 | 0.011 |
| Ethi+Latn | 138 | 0.008 | 0.008 |
| Arab+Ethi | 24 | 0.007 | 0.010 |
| Arab+Arab | 12 | 0.023 | 0.066 |
| Ethi+Ethi | 6 | 0.001 | 0.006 |
Figure 7: No caption. a — Pairwise transfer effects across African languages, measured by (a) F\(_1\)-BTS (accuracy) and (b) loss-BTS (calibration). Rows indicate target languages and columns indicate source languages used for co-training; cell values show the transfer score, with warmer colors indicating positive transfer and cooler colors indicating negative transfer. Color bars along the axes indicate high-level language-family groupings. Dashed boxes highlight selected within-family or same-script clusters where transfer is especially structured, such as Arabic varieties, Southern Bantu languages, Mande languages, and Fulah varieties. Both panels share a common symmetric color scale; note that F\(_1\)-BTS and loss-BTS are distinct quantities, so the same pair may gain accuracy in (a) while losing calibration in (b).
| ISO-3 | Language | Sentences | ISO-3 | Language | Sentences | ISO-3 | Language | Sentences |
|---|---|---|---|---|---|---|---|---|
| aar | Afar | 2,930 | aba | Abé | 6,902 | abi | Abidji | 10,218 |
| abn | Abua | 8,431 | acd | Gikyode | 7,102 | ach | Acholi | 8,626 |
| ada | Dangme | 11,002 | ade | Adele | 14,293 | adh | Jopadhola | 7,569 |
| adj | Adioukrou | 7,121 | aeb | Arabic, Tunisian | 12,985 | afr | Afrikaans | 45,990 |
| agq | Aghem | 2,942 | aha | Ahanta | 6,773 | ajg | Aja | 7,714 |
| aka | Akan | 2,948 | akp | Siwu | 7,118 | ald | Alladian | 13,794 |
| alz | Alur | 10,003 | amf | Hamer-Banne | 14,270 | amh | Amharic | 35,740 |
| ann | Obolo | 7,310 | anu | Anuak | 2,917 | anv | Denya | 7,126 |
| any | Anyin | 13,994 | ara | Arabic | 2,938 | arq | Arabic, Algerian | 3,693 |
| ary | Arabic, Moroccan | 24,745 | arz | Arabic, Egyptian | 25,958 | asa | Asu | 2,939 |
| asg | Cishingini | 6,884 | atg | Ivbie North-Okpela-Arhe | 7,135 | ati | Attié | 6,212 |
| avn | Avatime | 6,638 | avu | Avokaya | 7,080 | azo | Awing | 2,627 |
| bam | Bamanankan | 11,854 | bas | Basaa | 8,895 | bav | Vengo | 6,800 |
| bba | Baatonum | 9,452 | bbj | Ghomálá’ | 6,605 | bbk | Babanki | 5,191 |
| bbo | Konabéré | 10,168 | bci | Baoulé | 12,880 | bcn | Bali | 2,942 |
| bcw | Bana | 7,119 | bcy | Bacama | 336 | bdh | Baka | 6,750 |
| bds | Burunge | 2,940 | bem | Bemba | 41,323 | beq | Beembe | 6,638 |
| bex | Jur Modo | 7,389 | bez | Bena | 2,946 | bfa | Bari | 2,936 |
| bfd | Bafut | 7,129 | bfo | Birifor, Malba | 6,824 | bib | Bisa | 7,133 |
| bim | Bimoba | 8,827 | bin | Edo | 9,811 | biv | Birifor, Southern | 7,124 |
| bjv | Bedjond | 7,569 | bkv | Bekwarra | 14,291 | bky | Bokyi | 2,929 |
| blh | Kuwaa | 10,202 | bmo | Chrambo | 2,940 | bmq | Bomu | 14,293 |
| bmv | Bum | 6,633 | bom | Berom | 6,676 | bov | Tuwuli | 7,117 |
| box | Buamu | 7,103 | bqc | Boko | 8,703 | bqj | Bandial | 6,635 |
| bqp | Bisã | 14,290 | bsc | Oniyan | 6,636 | bsp | Baga Sitemu | 6,332 |
| bsq | Bassa | 4,803 | bss | Akoose | 7,119 | bst | Basketo | 1,435 |
| btt | Bete-Bendi | 14,286 | bud | Ntcham | 6,825 | bum | Bulu | 10,172 |
| bun | Sherbro | 333 | bus | Bokobaru | 7,213 | buy | Bullom So | 668 |
| bwq | Bobo Madaré, Southern | 14,292 | bwr | Bura-Pabir | 2,929 | bwu | Buli | 7,114 |
| bxk | Bukusu | 2,904 | byf | Bete | 2,681 | byv | Medumba | 3,902 |
| bza | Bandi | 2,942 | bzw | Basa | 2,939 | cce | Chopi | 9,805 |
| cgg | Chiga | 5,262 | chw | Chuwabu | 9,742 | cjk | Chokwe | 21,845 |
| cko | Anufo | 6,786 | cme | Cerma | 7,105 | cop | Coptic | 9,302 |
| cou | Wamey | 6,042 | cri | Sãotomense | 2,697 | crs | Seychelles French Creole | 13,500 |
| csk | Jola-Kasa | 7,127 | cwe | Kwere | 7,120 | cwt | Kuwaataay | 14,288 |
| daa | Dangaléat | 7,104 | daf | Dan | 13,673 | dag | Dagbani | 9,656 |
| dav | Dawida | 2,944 | dbq | Daba | 10,081 | ddn | Dendi | 2,238 |
| dga | Dagaare, Southern | 6,991 | dgd | Dagaari Dioula | 2,942 | dgi | Dagara, Northern | 6,829 |
| dhm | Dhimba | 6,634 | dib | Dinka, South Central | 1,175 | did | Didinga | 5,898 |
| dig | Chidigo | 6,828 | dik | Dinka, Southwestern | 17,349 | din | Dinka | 340 |
| dip | Dinka, Northeastern | 6,051 | diu | Gciriku | 906 | dje | Zarma | 2,749 |
| dks | Dinka, Southeastern | 6,965 | dnj | Dan | 7,216 | dop | Lukpa | 15,230 |
| dos | Dogosé | 10,232 | dow | Doyayo | 6,643 | dsh | Daasanach | 6,019 |
| dts | Dogon, Toro So | 14,276 | dua | Duala | 8,520 | dug | Chiduruma | 6,824 |
| dur | Dii | 13,791 | dwr | Dawro | 7,582 | dyi | Sénoufo, Djimini | 7,105 |
| dyo | Jola-Fonyi | 3,800 | dyu | Jula | 20,860 | ebr | Tchaman | 2,939 |
| ebu | Kiembu | 2,915 | efi | Efik | 16,007 | ego | Eggon | 2,942 |
| eka | Ekajuk | 7,120 | eko | Koti | 4,636 | enb | Markweeta | 13,700 |
3.0pt
| ISO-3 | Language | Sentences | ISO-3 | Language | Sentences | ISO-3 | Language | Sentences |
|---|---|---|---|---|---|---|---|---|
| eto | Eton | 2,286 | etu | Ejagham | 6,635 | etx | Iten | 2,936 |
| ewe | Éwé | 42,255 | ewo | Ewondo | 6,829 | fak | Fang | 2,854 |
| fal | Fali, South | 14,292 | fan | Fang | 14,513 | ffm | Fulfulde, Maasina | 7,508 |
| fia | Nobiin | 345 | fip | Fipa | 2,944 | flr | Fuliiru | 2,924 |
| fon | Fon | 22,482 | fub | Fulfulde, Adamawa | 8,041 | fue | Fulfulde, Borgu | 6,615 |
| fuf | Pular | 7,355 | fuh | Fulfulde, Western Niger | 7,154 | ful | Fulah | 2,513 |
| fuq | Fulfulde, Central-Eastern Niger | 6,737 | fuv | Fulfulde, Nigerian | 16,142 | gaa | Ga | 16,520 |
| gax | Oromo, Borana-Arsi-Guji | 2,852 | gaz | Oromo, West Central | 37,860 | gbo | Grebo, Northern | 5,684 |
| gbr | Gbagyi | 6,617 | gde | Gude | 7,118 | gej | Gen | 10,257 |
| gid | Gidar | 6,628 | giz | Giziga | 7,947 | gjn | Gonja | 7,833 |
| gkn | Gokana | 9,191 | gkp | Kpelle, Guinea | 2,898 | gmv | Gamo | 7,987 |
| gna | Kaansa | 6,614 | gnd | Zulgo-Gemzek | 6,815 | gng | Ngangam | 7,106 |
| goa | Guro | 1,929 | gof | Gofa | 7,470 | gog | Gogo | 6,869 |
| gol | Gola | 2,941 | gqr | Gor | 7,129 | gso | Gbaya, Southwest | 7,168 |
| gud | Dida, Yocoboué | 6,629 | guk | Gumuz | 14,292 | gur | Farefare | 9,570 |
| guw | Gun | 14,422 | gux | Gourmanchéma | 7,617 | guz | Ekegusii | 6,546 |
| gvl | Gulay | 6,918 | gwr | Gwere | 6,046 | gya | Gbaya, Northwest | 8,038 |
| hae | Oromo, Eastern | 14,293 | hag | Hanga | 4,840 | har | Harari | 114 |
| hau | Hausa | 32,959 | hav | Havu | 11,013 | hay | Haya | 6,848 |
| hbb | Nya Huba | 2,939 | heh | Hehe | 7,169 | her | Herero | 10,708 |
| hgm | Hai|ǁom | 2,888 | hig | Kamwe | 3,696 | hna | Mina | 2,937 |
| ibb | Ibibio | 2,945 | ibo | Igbo | 35,306 | idu | Idoma | 8,803 |
| ife | Ifè | 14,528 | igb | Ebira | 2,942 | ige | Igede | 8,879 |
| igl | Igala | 2,947 | ijn | Kalabari | 2,932 | ikk | Ika | 8,221 |
| ikw | Ikwere | 7,128 | ilb | Ila | 13,584 | iqw | Ikwo | 6,631 |
| iri | Rigwe | 7,130 | irk | Iraqw | 14,292 | ish | Esan | 9,493 |
| iso | Isoko | 14,172 | iyx | Yaka | 853 | izr | Izere | 7,132 |
| izz | Izii | 8,063 | jbu | Jukun Takum | 14,292 | jgo | Ngomba | 2,939 |
| jib | Jibu | 2,933 | jit | Jita | 2,939 | jmc | Machame | 6,752 |
| kab | Kabyle | 36,057 | kam | Kamba | 24,505 | kao | Xaasongaxango | 14,292 |
| kbn | Kare | 2,939 | kbo | Keliko | 5,192 | kbp | Kabiyè | 26,519 |
| kbr | Kafa | 14,182 | kby | Kanuri, Manga | 2,502 | kcg | Tyap | 6,639 |
| kck | Kalanga | 7,593 | kdc | Kutu | 7,122 | kde | Makonde | 7,631 |
| kdh | Tem | 4,320 | kdi | Kumam | 7,020 | kdj | Ng’akarimojong | 6,717 |
| kdl | Tsikimba | 7,119 | kdn | Kunda | 2,924 | kea | Kabuverdianu | 21,334 |
| ken | Kenyang | 7,120 | keo | Kakwa | 2,363 | ker | Kera | 10,211 |
| kez | Kukele | 14,291 | khq | Songhay, Koyra Chiini | 10,230 | khy | Kele | 6,639 |
| kia | Kim | 8,214 | kik | Gikuyu | 28,959 | kin | Kinyarwanda | 53,340 |
| kiz | Kisi | 2,885 | kki | Kagulu | 7,132 | kkj | Kako | 7,127 |
| kln | Kalenjin | 2,871 | klu | Klao | 2,932 | kma | Konni | 6,042 |
| kmb | Kimbundu | 26,546 | kmy | Koma | 5,307 | knc | Kanuri, Yerwa | 10,991 |
| knf | Mankanya | 7,194 | kng | Koongo | 7,442 | knk | Kuranko | 6,818 |
| kno | Kono | 6,805 | kny | Kanyok | 5,261 | kon | Kongo | 2,948 |
| koo | Konzo | 11,112 | koq | Kota | 390 | kpz | Kupsapiiny | 3,741 |
| kqn | Kaonde | 13,753 | kqo | Krahn, Eastern | 13,800 | kqp | Kimré | 7,125 |
| kqs | Kissi, Northern | 6,629 | kqy | Koorete | 6,796 | kri | Krio | 10,861 |
| krs | Gbaya | 2,947 | krw | Krahn, Western | 2,932 | krx | Karon | 1,201 |
| ksb | Shambala | 6,627 | ksf | Bafia | 6,655 | ksp | Kabba | 6,383 |
3.0pt
| ISO-3 | Language | Sentences | ISO-3 | Language | Sentences | ISO-3 | Language | Sentences |
|---|---|---|---|---|---|---|---|---|
| kss | Kisi, Southern | 7,796 | ktb | Kambaata | 14,253 | ktj | Krumen, Plapo | 7,098 |
| ktu | Kituba | 34,578 | kua | Oshiwambo | 13,681 | kub | Kutep | 7,127 |
| kuj | Kuria | 6,616 | kus | Kusaal | 6,834 | kvj | Psikye | 6,573 |
| kwn | Kwangali | 12,072 | kwy | Kikongo | 20,333 | kxc | Konso | 14,289 |
| kyf | Kouya | 6,851 | kyq | Kenga | 7,512 | kzn | Kokola | 2,789 |
| kzr | Karang | 2,695 | lai | Lambya | 6,618 | laj | Lango | 7,356 |
| lam | Lamba | 8,235 | lap | Laka | 4,931 | las | Lama | 10,206 |
| ldi | Laari | 15,818 | lea | Lega-Shabunda | 4,750 | led | Lendu | 5,573 |
| lee | Lyélé | 7,109 | lef | Lelemi | 6,826 | leh | Lenje | 15,944 |
| lem | Nomaande | 7,128 | lgg | Lugbara | 8,842 | lgm | Lega-Mwenga | 6,631 |
| lia | Limba, West-Central | 6,832 | lik | Lika | 2,929 | lin | Lingala | 43,654 |
| lip | Sekpele | 7,126 | llb | Lolo | 14,888 | lln | Lele | 12,030 |
| lmd | Lumun | 2,656 | lmp | Limbum | 6,628 | lnl | Banda, South Central | 2,924 |
| lob | Lobi | 14,293 | log | Logo | 4,561 | lok | Loko | 10,225 |
| lol | Mongo-Nkundu | 13,799 | lom | Loma | 6,396 | loq | Lobala | 6,308 |
| lot | Otuho | 2,939 | loz | Lozi | 15,985 | lro | Laro | 2,939 |
| lsm | Saamya-Gwe | 6,815 | lth | Thur | 2,895 | lto | Olutsotso | 2,885 |
| lua | Luba-Kasai | 39,225 | lub | Luba-Katanga | 14,269 | luc | Aringa | 6,501 |
| lue | Luvale | 14,312 | lug | Ganda | 31,230 | lun | Lunda | 13,024 |
| luo | Dholuo | 32,868 | lwg | Oluwanga | 3,411 | lwo | Luwo | 7,107 |
| maf | Mafa | 6,642 | mas | Maasai | 8,400 | maw | Mampruli | 7,113 |
| mbu | Mbula-Bwazza | 2,945 | mck | Mbunda | 10,430 | mcn | Masana | 7,797 |
| mcp | Makaa | 7,356 | mcu | Mambila, Cameroon | 6,825 | mda | Mada | 7,128 |
| mdm | Mayogo | 2,932 | mdy | Male | 8,977 | men | Mende | 6,891 |
| meq | Merey | 7,120 | mer | Kimîîru | 4,678 | mev | Maan | 2,415 |
| mfe | Morisyen | 13,701 | mfg | Mogofin | 4,216 | mfh | Matal | 6,833 |
| mfi | Wandala | 7,118 | mfk | Mofu, North | 7,130 | mfq | Moba | 6,677 |
| mfz | Mabaan | 4,007 | mgc | Morokodo | 5,449 | mgh | Makhuwa-Meetto | 9,323 |
| mgo | Meta’ | 6,638 | mgq | Malila | 2,939 | mgr | Mambwe-Lungu | 10,507 |
| mgw | Matumbi | 1,636 | mhi | Ma’di | 3,722 | mhw | Mbukushu | 45 |
| mif | Mofu-Gudur | 7,103 | mkl | Mokole | 7,119 | mlg | Malagasy | 2,481 |
| mlr | Vame | 887 | mmy | Migaama | 2,710 | mnf | Mundani | 6,821 |
| mnk | Mandinka | 6,832 | mny | Manyawa | 15,521 | moa | Mwan | 7,127 |
| mor | Moro | 3,695 | mos | Moore | 30,772 | moy | Shekkacho | 2,945 |
| moz | Mukulu | 2,928 | mpe | Majang | 2,932 | mpg | Marba | 7,119 |
| mqb | Mbuko | 7,122 | msc | Maninka, Sankaran | 5,067 | mse | Musey | 13,796 |
| mua | Mundang | 13,796 | mug | Musgu | 15,112 | muh | Mündü | 14,539 |
| mur | Murle | 6,814 | muy | Muyang | 6,933 | mwe | Mwera | 2,946 |
| mwm | Sar | 7,955 | mwn | Nyamwanga | 8,332 | mws | Mwimbi-Muthambi | 877 |
| myb | Mbay | 7,124 | myk | Sénoufo, Mamara | 7,155 | myx | Masaaba | 9,251 |
| mzk | Mambila, Nigeria | 14,293 | mzm | Mumuye | 7,139 | mzw | Deg | 7,089 |
| naq | Khoekhoe | 9,663 | naw | Nawuri | 7,133 | nba | Nyemba | 10,339 |
| nbl | Ndebele | 13,722 | ncu | Chumburung | 7,081 | ndc | Ndau | 11,260 |
| nde | Ndebele | 11,979 | ndh | Ndali | 2,455 | ndi | Samba Leko | 13,793 |
| ndj | Ndamba | 7,135 | ndo | Ndonga | 13,803 | ndp | Kebu | 14,293 |
| ndv | Ndut | 2,728 | ndy | Luto | 10,245 | ndz | Ndogo | 7,130 |
| neb | Toura | 14,284 | nfr | Nafaanra | 14,290 | ngb | Ngbandi, Northern | 4,833 |
| ngc | Ngombe | 6,639 | ngl | Lomwe | 10,462 | ngn | Ngwo | 2,899 |
3.0pt
| ISO-3 | Language | Sentences | ISO-3 | Language | Sentences | ISO-3 | Language | Sentences |
|---|---|---|---|---|---|---|---|---|
| ngp | Ngulu | 7,137 | nhr | Naro | 7,113 | nhu | Noone | 7,111 |
| nih | Nyiha, Tanzania | 2,939 | nim | Nilamba | 7,137 | nin | Ninzo | 7,253 |
| niq | Nandi | 7,786 | niy | Ngiti | 6,638 | nka | Nkoya | 2,842 |
| nko | Nkonya | 7,107 | nla | Ngombale | 2,341 | nmz | Nawdm | 14,289 |
| nnb | Nande | 18,686 | nnh | Ngiemboon | 6,324 | nnq | Ngindo | 7,133 |
| nnw | Nuni, Southern | 14,526 | nqo | N’Ko | 12,823 | nse | Nsenga | 9,637 |
| nso | Sotho, Northern | 46,564 | ntr | Delo | 7,129 | nuj | Nyole | 7,658 |
| nus | Nuer | 11,861 | nwb | Nyabwa | 6,776 | nxd | Ngando | 6,644 |
| nya | Chichewa | 51,207 | nyb | Nyagbo | 2,940 | nyd | Olunyole | 2,942 |
| nyf | Kigiryama | 6,848 | nyk | Nyaneka | 11,885 | nym | Nyamwezi | 2,934 |
| nyn | Nyankore | 11,475 | nyo | Nyoro | 6,841 | nyu | Nyungwe | 10,624 |
| nyy | Nyakyusa-Ngonde | 9,740 | nza | Mbembe, Tigon | 5,188 | nzi | Nzema | 12,609 |
| odu | Odual | 2,938 | ogo | Khana | 8,464 | oke | Okpe | 8,523 |
| okr | Kirike | 2,941 | oku | Oku | 6,048 | old | Mochi | 13,340 |
| orm | Oromo | 2,792 | ozm | Koonzime | 6,799 | pbi | Parkwa | 14,279 |
| pcm | Pidgin, Nigerian | 13,317 | pem | Phende | 6,198 | pfe | Pere | 14,395 |
| phm | Phimbi | 15,399 | pkb | Kipfokomu | 6,828 | pko | Pökoot | 2,783 |
| plt | Malagasy, Merina | 27,852 | pny | Pinyin | 13,238 | pov | Guinea-Bissau Creole | 7,251 |
| poy | Pogolo | 6,644 | rag | Lulogooli | 2,900 | rcf | Réunion French Creole | 13,722 |
| rel | Rendille | 6,056 | rif | Tarifit | 2,941 | rim | Nyaturu | 7,137 |
| rnd | Ruund | 7,481 | rng | Ronga | 8,976 | rub | Gungu | 6,636 |
| ruf | Luguru | 14,533 | run | Rundi | 44,213 | rwk | Rwa | 2,633 |
| sag | Sango | 36,544 | saq | Samburu | 2,946 | sba | Ngambay | 8,083 |
| sbd | Samo, Southern | 6,654 | sbp | Sangu | 2,939 | sbs | Kuhane | 2,540 |
| sby | Soli | 2,407 | sef | Sénoufo, Cebaara | 2,935 | seh | Sena | 19,145 |
| ses | Songhay, Koyraboro Senni | 7,372 | sev | Sénoufo, Nyarafolo | 2,930 | sfw | Esahie | 6,359 |
| sgc | Kipsigis | 11,612 | sgw | Sebat Bet Gurage | 6,828 | shi | Tachelhit | 8,390 |
| shj | Shatt | 766 | shk | Shilluk | 5,943 | shr | Shi | 14,113 |
| shu | Arabic, Chadian | 6,626 | sid | Sidaama | 9,955 | sig | Paasaal | 7,124 |
| sil | Sisaala, Tumulung | 6,827 | skg | Malagasy, Sakalava | 15,615 | sld | Sissala | 14,290 |
| sna | Shona | 45,067 | snf | Noon | 6,634 | sng | Sanga | 2,913 |
| snw | Selee | 6,863 | soe | Ohendo | 717 | som | Somali | 27,870 |
| sop | Songe | 11,004 | sor | Soumraye | 1,083 | sot | Sotho, Southern | 13,079 |
| soy | Miyobe | 7,118 | spp | Sénoufo, Supyire | 7,114 | spy | Sabaot | 14,640 |
| srr | Serer-Sine | 3,905 | ssw | Swati | 27,224 | suk | Sukuma | 7,344 |
| sur | Mwaghavul | 14,293 | sus | Susu | 8,885 | swa | Swahili | 2,918 |
| swb | Comorian, Maore | 237 | swc | Swahili, Congo | 28,449 | swh | Swahili | 37,548 |
| swk | Sena, Malawi | 5,962 | sxb | Suba | 7,121 | tap | Taabwa | 10,437 |
| taq | Tamasheq | 15,981 | tbz | Ditammari | 5,589 | tcc | Datooga | 7,128 |
| tcd | Tafi | 2,947 | tdx | Malagasy, Tandroy-Mahafaly | 8,366 | ted | Krumen, Tepo | 6,861 |
| tem | Themne | 6,840 | teo | Ateso | 8,489 | tex | Tennet | 2,942 |
| tgw | Sénoufo, Tagwana | 2,934 | thk | Kitharaka | 6,820 | thv | Tamahaq, Tahaggart | 1,946 |
| tig | Tigré | 3,105 | tik | Tikar | 14,289 | tir | Tigrigna | 36,754 |
| tiv | Tiv | 13,164 | tke | Takwane | 6,888 | tlj | Talinga-Bwisi | 7,117 |
| tll | Tetela | 13,894 | tmc | Tumak | 10,258 | tnr | Ménik | 9,951 |
| tod | Toma | 7,199 | tog | Tonga | 11,237 | toh | Tonga | 9,554 |
| toi | Tonga | 15,080 | tpm | Tampulma | 8,553 | tsc | Tswa | 12,065 |
| tsn | Setswana | 35,384 | tso | Tsonga | 38,043 | tsw | Tsishingini | 7,125 |
3.0pt
| ISO-3 | Language | Sentences | ISO-3 | Language | Sentences | ISO-3 | Language | Sentences |
|---|---|---|---|---|---|---|---|---|
| ttj | Tooro | 9,664 | ttq | Tamajaq, Tawallammat | 6,624 | ttr | Tera | 2,946 |
| tui | Tupuri | 7,698 | tul | Tula | 5,200 | tum | Tumbuka | 34,918 |
| tuv | Turkana | 2,383 | tvu | Tunen | 2,942 | twi | Twi | 43,976 |
| twx | Tewe | 14,947 | tzm | Tamazight, Central Atlas | 2,472 | udu | Uduk | 4,024 |
| umb | Umbundu | 32,657 | urh | Urhobo | 10,605 | uth | ut-Hun | 5,189 |
| vag | Vagla | 7,010 | vai | Vai | 2,940 | ven | Venda | 17,236 |
| vid | Vidunda | 7,117 | vif | Vili | 2,942 | vmk | Makhuwa-Shirima | 4,231 |
| vmw | Makhuwa | 11,968 | vun | Vunjo | 7,131 | vut | Vute | 7,120 |
| wal | Wolaytta | 12,310 | wbi | Vwanji | 2,937 | wec | Wè Western | 2,691 |
| wes | Pidgin, Cameroon | 8,959 | wib | Toussian, Southern | 6,158 | wlx | Wali | 14,628 |
| wmw | Mwani | 7,125 | wob | Wè Northern | 14,529 | wol | Wolof | 17,637 |
| won | Wongo | 512 | wwa | Waama | 3,694 | xan | Xamtanga | 1,474 |
| xed | Hdi | 7,103 | xho | Xhosa | 37,847 | xmv | Malagasy, Antankarana | 15,974 |
| xnz | Mattokki | 2,939 | xog | Soga | 7,860 | xon | Konkomba | 7,511 |
| xpe | Kpelle, Liberia | 2,864 | xrb | Karaboro, Eastern | 6,933 | xsm | Kasem | 7,519 |
| xtc | Katcha-Kadugli-Miri | 2,946 | xuo | Kuo | 7,113 | yal | Yalunka | 7,086 |
| yam | Yamba | 7,124 | yao | Yao | 11,100 | yas | Nugunu | 2,168 |
| yat | Yambeta | 6,027 | yaz | Lokaa | 10,258 | yba | Yala | 2,744 |
| ybb | Yemba | 6,301 | yom | Kiyombe | 9,055 | yor | Yoruba | 48,191 |
| yre | Yaouré | 7,668 | zaj | Zaramo | 2,937 | zdj | Comorian, Ngazidja | 4,106 |
| zga | Kinga | 2,941 | zgh | Tamazight, Standard Moroccan | 5,817 | ziw | Zigula | 7,137 |
| zne | Zande | 11,952 | zul | Zulu | 37,885 |
3.0pt
https://github.com/cambridgeltl/mirror-bert↩︎