March 30, 2026
The design of Large Language Models (LLMs) and generative artificial intelligence (GenAI) has been shown to be “unfair” to less-spoken languages [1] and to deepen the digital language divide [2]. Critical sociolinguistic work has also argued that these technologies are not only made possible by prior socio-historical processes of linguistic standardisation, often grounded in European nationalist and colonial projects [3], but also exacerbate epistemologies of language as “monolithic, monolingual, syntactically standardized systems of meaning” [4]. In our paper, we draw on earlier work on the intersections of technology and language policy [5] and bring our respective expertise in critical sociolinguistics and computational linguistics to bear on an interrogation of these arguments. We take two different complexes of non-standard linguistic varieties in our respective repertoires–South Tyrolean dialects, which are widely used in informal communication in South Tyrol, Italy [6], as well as varieties of Kurdish–as starting points to an interdisciplinary exploration of the intersections between GenAI and linguistic variation and standardisation. We discuss both how LLMs can be made to deal with non-standard language from a technical perspective, and whether, when or how this can contribute to “democratic and decolonial digital and machine learning strategies” [3], which has direct policy implications.
The rapid advancement of Large Language Models (LLMs) and Generative Artificial Intelligence (GenAI) has transformed digital communication, yet these technologies systematically privilege standardized and, in computational linguistic terms, high-resource languages while marginalizing millions of speakers who communicate primarily through non-standard varieties and dialects. Recent scholarship demonstrates that LLM design is fundamentally “unfair” to less-spoken languages [1] and deepens the digital language divide [2]. Critical sociolinguistic work argues that these technologies not only emerge from historical processes of linguistic standardization rooted in colonial and nationalist projects [3], but also reinforce epistemologies of language as “monolithic, monolingual, syntactically standardized systems of meaning” [4]. When AI systems fail to process non-standard varieties, they exclude speakers from full participation in digital citizenship. Further, technical decisions made in the architecture of models impose a “tokenization tax” that increases costs and degrades performance for languages less represented in the input data to LLMs.
This paper brings together critical sociolinguistics and computational linguistics to examine these issues through two complementary case studies: South Tyrolean German and Kurdish varieties. South Tyrolean dialect comprises non-standard forms of German widely used in informal communication in South Tyrol, Italy [6], yet remains largely absent from LLM training data despite being the primary medium of everyday interaction. Kurdish represents a dialect continuum spoken by over 40 million people across multiple nation-states, characterized by orthographic diversity, systematic political suppression, and acute digital underrepresentation. While Central Kurdish (Sorani) and Northern Kurdish (Kurmanji) have achieved modest computational resources, varieties such as Southern Kurdish, Hawrami, and Zazaki remain almost entirely invisible to language technology. Through these cases, we analyze the technical barriers preventing LLMs from processing non-standard varieties, examine how evaluation frameworks reproduce standardization biases, and explore policy implications for diverse actors—from Big Tech and states to civil society and academia.
Building on earlier work examining technology and language policy intersections [5], we argue that addressing LLM marginalization of non-standard varieties requires more than technical fixes; it demands a reorientation toward “democratic and decolonial digital and machine learning strategies” [3] that center linguistic agency and treat variation as constitutive of human communication rather than noise to eliminate. This has direct policy implications, from requiring “dialect gap” reporting to ensuring community data sovereignty. Ultimately, whether LLMs should “handle” non-standard varieties is not primarily a technical question but a fundamentally political one, with implications for digital sovereignty, cultural preservation, and social justice in an age where the choice between enforcing standardization or embracing variation will shape the vitality of the world’s linguistic diversity in digital spaces. In the following sections, we will first address sociolinguistic (section 2) and computational linguistic perspectives on standard and non-standard language and language technologies (section 3), before giving an overview of the policy landscape in relation to linguistic variation and LLMs (section 4). We will proceed by examining the cases of South Tyrolean dialects (section [case95study95tyrolean]) and varieties of Kurdish (section 6) and end with overarching conclusions and policy implications.
The concept of standard language, as well as processes of linguistic standardisation, have been objects of sociolinguistic investigation since the inception of the field itself [7], [8]. The impetus for this type of work largely came from concurrent processes of standardization in newly independent states, often postcolonial realities. As such, it is closely linked to the beginnings of language planning and policy as a subfield of applied linguistics. As characteristics of standard languages, [9] lists their validity across regional borders, their being considered as the ‘best’ language within their realm of validity, and their being codified in norms. As [10], [11] notes, it is their institutional enforcement that distinguishes standard languages from other types of linguistic norms, both from “always emergent, variable, and never ‘fixed’ conventions” [11] and from language standards that go beyond such conventions in that they become morally imperative, but are not quite standard languages as they are not institutionally enforced.
Processes of standardisation usually involve some degree of reduction in linguistic variability [10], [12]. As such, these processes can also be understood as linguistic hierarchisations [13]–[15], whereby specific sets of linguistic features, as well as entire varieties conceived as bundles of such features, are imbued with legitimacy for public use, while others are considered acceptable only in private [14] . As [15] argues, this legitimacy and authority of linguistic forms stem the two ideological complexes of anonymity and authenticity that go back to rationalist and romantic philosophies of the 17th and 18th centuries.
It is commonly recognised that “[s]tandardization is […] best approached as an ideological phenomenon” [15] . Processes of linguistic standardisation have tended to be part and parcel of processes of nation-building [11], [14], [16] and the concomitant creation of an apparently neutral public sphere, as well as of colonialism, and are quintessentially modernist projects [14]. Standard language and non-standard language – be it referred to as dialect, patois, etc. – are constructed in opposition to one another, whereby only standard language indexes modernist values like progress and development, and is associated with the future [15]. It is easy to see how, within such standardisation regimes, the majority of the world’s languages that do not have standardised forms [17] can become constructed as ‘lacking’ and ‘backwards’ [15], and even more so the myriad of primarily oral language practices [18]. In fact, the epistemologies of language underlying standard language ideologies are those of language “as an autonomous and unitary system whose main function is the effective and precise transmission of information” [11], disregarding a view of language as situated and embodied social practice [4], [14].
It has been argued that the development of LLMs has been based both on the epistemological notions underlying standardising regimes and on the effects of these regimes on language practices: According to [4], LLMs “build on a foundation of prior technological and sociopolitical conditions including phonetic spelling, standardized print literacy, and linguistic nationalism” [4]. It is thus the diffusion of notions of standard language along with universal education and mass literacy that have produced a quantity of text homogeneous enough to be probabilistically modellable – and these models, in turn, provide the basis for the generation of statistically likely sequences of tokens by generative AI [4].
While linguistic standardisation is usually examined within the bounds of a single nation, a different perspective is necessary for investigating the development and effects of language technologies like LLMs and generative AI. As [19] notes, in contrast with the modernist projects of nation-building and linguistic standardisation, these technologies do not aim “to homogenize language in order to create a linguistically homogenous national population” but instead produce a global digital public and are motivated, in most cases, by commercial interests. Previous work on the intersections of technology and language policy in relation to the development of the internet has shown how “technological advances and breakthroughs occur in particular ideological and cultural spaces, and the shape of those technological advances bears the imprint of those cultural and ideological norms” [5]. Dividing the de facto language policy of the internet into four distinct periods, [5] argued that we are currently witnessing the period of idiolingualism, characterised by mass linguistic customization. It is to be expected that this type of customization or personalization will impact on the linguistic direction that LLMs will take, and that user needs – as well as their power of consumption – will contribute to steering their development [19], [20]
In this manner, language technologies may reinforce and potentially also reconfigure linguistic hierarchisations [19], [21]. As we will also show in our next section, the hierarchies originating from the interplay between linguistic standardisation and colonialism are already being exacerbated, with LLMs further contributing to the dominance of English and of other European-derived standard languages [19], and of Mandarin [22]. At the same time, however, the jury is still out on the effects that this technology will have on variation within what is commonly constructed as one language. For instance, the fact that the acceptance of standard languages as ‘best’ language tends to be more wide-spread than their use [13], [15] might mean that it might not be those forms of language that LLMs will reinforce, if other forms end up being used more frequently – analogous to destandardisation tendencies that have been identified in different European contexts for some time now [9]. It thus becomes a highly relevant question how LLMs and generative AI will impact on hierarchisations of language practices and their associated types of speakers.
Underlying those generative AI systems that generate text, Large Language Models (LLMs) have demonstrated remarkable capabilities in what computational linguistics refers to as high-resource linguistic environments [23], especially English. However, their deployment across the global linguistic landscape remains deeply asymmetrical. In computational linguistic terms, the “digital divide” between languages has translated into lesser-used languages in the digital sphere becoming classed as Under-Resourced Languages (URLs), lacking representative training data and thus language resources. At the same time, these languages are impaired by Anglo-centric architectural biases of LLMs which negatively affect performance [24]. While this challenge is substantial for national or regional minority languages, like Irish or Basque, non-standardized dialects and varieties such as South Tyrolean or Kurdish varieties have been barely considered when evaluating factors such as LLM performance, even though LLMs can work to some degree in such languages [25]. Applying LLMs to these languages is often complex due to issues such as the lack of formal orthographies, diglossic tension with standard versions (like Standard German) [26] and they are even frequently deliberately excluded from the massive web-scrapes that form LLM training sets [27]. Consequently, these languages face a double marginalization where both the data scarcity and the structural assumptions of modern NLP fail to capture their unique phonetic, syntactic, and cultural nuances. We see three main areas where current research on LLMs can be improved for non-standard languages: The architecture of LLMs, especially around the tokenization of the input text, the handling of morphological complexity and orthographic variation, and the creation of relevant benchmarks.
The technical infrastructure of language technologies significantly impacts dialects and determines whether a variety is even visible to digital tools. One issue in this regard is the assignment of ISO-639 codes, which acts as a primary gatekeeper; for instance, German dialects like Upper Saxon or Bavarian have assigned ISO codes, which allows them to be catalogued in major Natural Language Processing (NLP) resources like OPUS and the Virtual Language Observatory. Without such codes, whose colonial history [3] have elaborated on, a variety is effectively invisible to language technology, making it nearly impossible to track its representation or performance in AI models. Moreover, the transcription of some dialects may make use of diacritics or other writing methods not supported by the Unicode standard, and as such cannot be processed. The graphemic representation of any language must therefore either remain within the existing coding, or the respective signs must be added to the Unicode standard.
To see why LLMs struggle with non-standard languages like South Tyrolean or Kurdish, we must also look at the representation system that stands between human text and the machine, namely the tokenizer. Modern AI models do not read text as humans do, but instead convert the input text into a sequence of numbers that can be processed by a neural network. The process of segmenting the text is called tokenization and most current LLMs use methods based on Byte Pair Encoding [28]. Instead of breaking text into whole words (which would create a vocabulary too large for the computer to manage) or individual characters (which are too small to carry much meaning), BPE identifies “subword” units based on how frequently they appear in a training dataset. As such models break words based on their statistical frequency, they will have more tokens assigned to languages that occur more frequently in the corpus. In this way, words for well-resourced languages will often be broken into fewer tokens and in ways that are more meaningful and related to the morphology of the language. For example, the English word ‘tokenization’ is divided into two tokens by the widely used BERT model1 [29], namely ‘token+ization’, separating the root and the suffix in a morphologically sound way. In contrast, the Irish word ‘ionchomharthú’ (meaning ‘tokenization’) is divided into 5 tokens, ‘ion+cho+m+hart+hú’, and this tokenization not only does not bare any resemblance to the word’s morphological components (ion+chomh+arth+ú), but even breaks the word across digraphs such as ‘mh’ and ‘th’2.
When a tokenizer encounters an under-resourced language or non-standard language, its training hasn’t resulted in any specialized subword entries for that language. Instead, it produces a sequence of tokens formed from the subwords that it already has in its vocabulary. This is computationally inefficient, as a single dialectal word might be shattered into 5 or 6 tiny fragments (e.g., individual characters or meaningless byte-strings), whereas the English equivalent would be a single token. This computational inefficiency results in extra costs, e.g. for using more electricity to process a query in these languages. Providers often pass on this cost to users, charging them per token, which causes what is referred to as a "tokenization tax"; so a speaker of Kurdish or South Tyrolean literally pays more when they prompt generative AI in these varieties3 to process the same amount of information given in English. Performance, in terms of the model’s ability, will also naturally be degraded as the model is working with small tokens without specific meaning, which carry less input than large subwords or whole words. Finally, most AI models have a context window [30], which is a hard limit on how many tokens it can consider at one time. Because dialects require more tokens to express the same idea, they fill up the model’s memory faster, leading to poorer reasoning and shorter possible conversations. In this light, the “tokenization tax” is not merely a technical detail [31]; it is a form of algorithmic discrimination [32] that systematically increases the cost and decreases the quality of AI services for marginalized linguistic communities.
A second major issue that further compounds the challenge of tokenization is that many under-resourced languages have high morphological complexity and non-standard languages have substantial orthographic variation. English and Chinese, languages spoken by many LLM developers, are relatively “morphologically poor” (such as measured by [33]), having fewer morphemes and simpler morphotactics. Tokenization methods are inherently biased towards morphologically poor languages, as languages with a lower Type-To-Token (TTR) ratio [34] can be represented with a smaller vocabulary of subwords. On the other hand, Kurdish for example is a fusional language, where a single word can be built by combining morphemes for functions such as tense, person or negation. As such, a whole sentence in English may be compressed into a single complex Kurdish word. This challenge is intensified by the presence of allomorphs: Because Kurdish grammar is highly sensitive to the sounds surrounding a prefix or suffix, a single grammatical marker (like a plural or a tense indicator) might change its spelling or sound depending on the verb it attaches to.
For languages without a unified orthography, multiple spellings may exist based on local sub-dialects. Moreover, much of what LLMs see of such languages is extracted from informal contexts such as social media, which often contain errors due to carelessness (i.e., typos) as well as linguistic variation. Further, for a language such as South Tyrolean, there is a strong linguistic pull towards the associated standard language (i.e., Modern High German) not only for the speakers but also for the model, which will have substantial training on the standard language. If AI tools are used for regional governance or service delivery, their inability to handle non-standard spelling can lead to exclusionary bias. Citizens who write in their native dialect may find themselves misunderstood or ignored by automated systems that were optimized for a standardized “prestige” language.
The final barrier to linguistic equity is how we measure LLM success. In AI development, “what gets measured gets built”4 [35]; however, the current tools for evaluating model performance are fundamentally ill-suited for non-standard and under-resourced languages. Measurement of an LLM’s general knowledge is achieved through benchmarks such as MMLU [36], which contains questions from standardised textbooks and similar sources. The MMLU benchmark is highly US-centric, including topics such as US History, US Law and US Accounting, with approximately 28% of the questions requiring specific knowledge of Western cultures and a staggering 84.9% of geographic questions focus exclusively on North America or Europe [25]. To address this, Global MMLU was introduced, which covers 42 languages, including highly under-resourced languages such as Nyanja or Telugu and explicitly marks questions as culturally sensitive. This avoids a distorted ranking, where a model might appear highly capable in a target language simply because it has memorised Western facts while failing to grasp the specific cultural, legal, or social nuances relevant to actual speakers of that language.
This bias is largely due to the fact that most benchmarks are created by translation of English benchmarks into other languages, and while this has been shown to correlate with human judgements [37], benchmarks are even worse for lower-resourced languages. This creates a translation pivot trap, where models are optimized to perform well on translated English concepts rather than achieving true native-level reasoning. To correct this bias, high-quality, culturally grounded benchmarks need to be developed, which is an immense financial and logistical undertaking. For example, the development of MMLU-ProX [24] involved a rigorous expert-review process to ensure cultural relevance, with development costs approaching $80,000 at market rates. Furthermore, the quality of benchmarks created by translation varies significantly with STEM-related tasks exhibiting strong correlation with human judgements (0.70-0.85), while other tasks, such as question answering, have very poor correlation (0.11-0.30) [38].
The path forward requires moving away from translated benchmarks toward “culturally and linguistically tailored benchmarks” [38]. A recent example is IRLBench [39], a benchmark for the Irish language derived from the Irish Leaving Certificate exams. Recently, DialectBench [40] has introduced a benchmark covering 281 dialects across 10 different tasks, providing the first benchmark that covers a wide variety of non-standard languages. However, South Tyrolean is absent from this benchmark and the support for Kurdish dialects is only partial. Further, there are specific aspects of relevance to non-standard languages that are excluded from benchmarks derived from English that are of relevance to speakers of the community. Firstly, being able to distinguish between dialects and measure whether a text is truly in the dialect or is just outputting standard language with a few dialectal words thrown in. This has led researchers to propose a Dialect Fidelity Score [41] to measure this. Secondly, more culturally relevant questions, for example, explaining the moral of a specific proverb, would be relevant, where success requires an understanding of the cultural metaphors that do not exist in the model’s English-centric training data. Finally, as non-standard languages are primarily used orally or in informal digital spaces, tasks in these benchmarks should reflect this bias. For example, VoxLect [42] uses speech foundation models to evaluate how well AI understands regional accents and phonetic variations that are never captured in written text. Similarly, using “noisy” data from social media, such as in the NorDial benchmark for Norwegian dialects [43] helps determining whether models can handle inconsistent orthography and non-standard spelling without crashing or defaulting to English or standard languages.
Considering the caveats of most current benchmarks, it is all the more alarming that even the best-performing models still show a persistent gap [44] and produce significantly worse results in languages other than English. For non-standard varieties like South Tyrolean or Kurdish, where formal academic benchmarks do not exist or are fragmentary, the “performance cliff” is likely even steeper. Without a policy-driven investment in local, non-translated evaluation data, speakers of non-standard varieties will not be able to make use of generative AI in these varieties.
In this section, we turn to intersections between linguistic variation, language policy and LLMs and show how the way in which LLMs work is based on earlier language policy, and which types of policy actors respond in which kinds of ways to this technology and its development.
Non-standard languages, in computational linguistic terms, are often referred to as “severely” or “acutely” under-resourced” [45], [46]. Traditionally, language documentation has been framed as an intervention intended to “save” endangered languages by capturing their lexico-grammatical structures as data before they are lost [18]. This model typically involves linguists creating resources, such as orthographies, lexicons, and grammars, to support the production of pedagogical materials for formal language programs, but also increasingly to develop artificial intelligence systems [47]. This is often founded on a deep-seated ideological bias of ‘scriptism’ [48], which treats written language as the primary, superior, or only “proper” form of language, often reducing speech to a mere derivative of text. In the context of language technology and documentation, scriptism manifests as the insistence on standardizing orthographies and prioritizing textual data as a prerequisite for technological support. Scriptism has led LLMs to ignore non-standardised languages by prioritising the creation of massive text-based datasets, which effectively excludes primarily oral or non-standard speech from the “resource horizon” of modern AI [40]. By treating standardized writing as the only suitable data and by developing systems that disregard linguistic variation, LLM technologies overlook the linguistic practices used by many in the world to communicate.
Beyond the bias of scriptism, the evolution of large language models is fundamentally linked to the hegemony of English as the global lingua franca of digital infrastructure and academic research. AI development has seen English as the default language [49] and as such, computational benchmarks treat English as a universal means of expression and the primary pivot language for multilingual capabilities. The reliance on English, thus, not only affects the amount of data available but also imposes Anglo-centric semantic structures and pragmatic norms onto other languages. Further, educational policies which are rooted in nationalistic and colonial projects [3] have further exacerbated this by ensuring that the majority of high-quality digitized texts, such as textbooks, academic papers and documentation, are produced in standardized language. This creates a feedback loop, where LLMs trained on these corpora understand these languages and registers as the means for intellectual discussions and thus generate descriptions in these registers when asked to explain complex and challenging topics. This leads to LLMs considering content that falls outside of these languages and registers to be of lesser intellectual value and less prestigious [50].
The rapid proliferation of LLMs and Generative AI has led to questions about how language policy can support the development of these technologies in a manner that supports language equality. By analyzing the motivations and limitations of diverse policy actors from Big Tech and state entities to civil society and academic institutions, we argue for a shift towards democratic machine learning strategies that recognize linguistic variation not as statistical noise, but as a vital component of human communication and digital citizenship.
The motivation for big technology companies like Meta, Google or OpenAI to move beyond standardized English is rarely purely altruistic but is governed by a tension between economic scalability and socio-technical responsibility. Big Tech actors occupy a contradictory space as both the primary enforcers of linguistic standardization and the only entities with the compute power to technically reduce the dialect gap. The dominant logic of the large scale favours standardisation to minimise computational costs; however, increasingly, performance on English cannot easily be improved and, as such, under-resourced and non-standard languages have become strategic for improving overall model performance. Similarly, high-resource language markets, such as English, Spanish or Mandarin, are saturated, while the “next billion web users” [22] speak a diverse range of languages. Development in this logic is likely to remain mostly extractive, unless governed by policies that support linguistic equality. A key technical measure could be the measurement and reporting of a dialect gap, which measures the reduction in performance in dialects versus English and by requiring this to be explicitly reported or even supported by means of Corporate Social Responsibility (CSR) credits. In this manner, the paradigm could be shifted towards more explicit support for non-standardized languages.
State actors have emerged as critical counter-weights to the English-centricity of commercial GenAI, reframing linguistic diversity as a pillar of digital sovereignty. Initiatives such as OpenEuroLLM5 support development in particular in EU official languages and the democratic participation of all its citizens. By leveraging public supercomputing infrastructure, such as the EuroHPC network6, states can subsidize the high computational cost of training models on acutely under-resourced dialect data. Furthermore, the implementation of the EU AI Act in 2026 provides a regulatory framework that requires high-risk AI systems to be transparent and non-discriminatory. This provides a legal hook: if an AI used in Italian public administration fails to understand a South Tyrolean citizen, it may be deemed a violation of fundamental rights to non-discrimination. While most modern states now frame linguistic diversity as a public good, historically, states have implemented nationalistic policies that marginalize non-standard varieties. This legacy has both led to the data voids that create issues with current LLMs, as well as reduced the trust among speakers of public initiatives. As such, state-led AI initiatives should avoid replicating the discriminatory practices of the past, e.g. by adopting a policy framework of linguistic agency, centered on the principle of “nothing about us without us.” A democratic strategy [3] requires that speakers of non-standard varieties retain data sovereignty over their linguistic repertoires.
Community and civil society organisations (CSOs) serve as the essential connective tissue between the technical requirements of LLM development and the lived reality of speech communities. For non-standard languages, like South Tyrolean dialect or varieties of Kurdish, CSOs can transition from being passive subjects of study to active data stewards and algorithmic auditors. Unlike Big Tech’s extractive scraping, CSOs can run citizen science initiatives [51] (e.g., using platforms like Mozilla Common Voice [52]) to collect authentic speech and text with explicit community consent. Furthermore, CSOs can serve as algorithmic auditors, performing socio-pragmatic red-teaming to identify where models fail to respect local norms or inadvertently enforce standardization. Similarly, academic institutions and national language institutes serve as the primary bridge between technical innovation and socio-historical depth. While Big Tech prioritizes computational scale, universities provide the sociolinguistic granularity necessary to prevent the erasure of non-standard varieties. Furthermore, by fostering interdisciplinary collaboration between Natural Language Processing engineers and sociolinguists, academics ensure that the models do not merely replicate standardized norms, but reflect a linguistic reality grounded in actual usage.
South Tyrol is the northernmost Italian province and is known for its contested history since its annexation to Italy in 1919, its autonomy provisions and its institutional multilingualism. German is one of the official languages in the Province – legally on par with Italian – but predominantly it is not Standard German that is spoken in social life in the province, but an ensemble of non-standard forms of German [53]. Research has stressed the social relevance of these local dialects [54], which are also increasingly being used in informal writing, particularly in digital contexts such as on social networks or on WhatsApp (see e.g. [6], [55]). Standard German, in contrast, has been shown to mostly be used as a spoken language in a school context [56], and while South Tyrolean German speakers seem to orient to a standard from Germany as the norm with the highest prestige, they have been shown to consider this standard as somewhat ‘foreign’ to them [57]. At the same time, however, most formal writing still takes place in standard German.
When evaluating the resourcefulness, in terms of available language resources and technologies, of South Tyrolean dialects, one first runs up against the problem that South Tyrolean does not have a specific ISO-639 code, which, as previously discussed, means it is not catalogued in NLP resources. As Table 1 shows, there are only ISO-codes for broader groups of German dialects, such as Upper Saxon, Allemanic, Swabian or Bavarian, which makes only language resources with these codes trackable. South Tyrolean dialects are thereby included within the broader group of Bavarian dialects, spoken through most of Austria (with the exception of Vorarlberg in the west) and most of Bavaria. While the Table shows that there are some language resources for Bavarian - especially in comparison to many other German dialect groups represented - it is impossible to see at first glance whether this actually also represents South Tyrolean varieties.
| Name | ISO Code | Speakers | OPUS Words | VLO Resources | |
|---|---|---|---|---|---|
| Upper Saxon | sxu | 2,000,000 | 0 | 0 | |
| Kölsch | ksh | 250,000 | 6,940 | 8 | |
| Palatine | pfl | 400,000 | 122 | 0 | |
| East Franconian | vmf | 4,900,000 | 0 | 0 | |
| Allemanic | gsw | 7,162,000 | 3,030 | 1,282 | |
| Swabian | swg | 820,000 | 203 | 8 | |
| Walser | wae | 22,780 | 0 | 0 | |
| Bavarian | bar | 15,000,000 | 156,237 | 124 |
The same then holds for evaluating the availability of pre-trained language models or benchmarks. While, as [25] note, LLMs might work to some degree in dialects - and in fact several Generative AI models responded affirmative when we asked "Versteasch du mi?" [Do you understand me? Standard German Verstehst du mich?] - dialects without a specific ISO code are so far not officially supported, and LLMs performance is not evaluated against them. However, the same is also true for many non-standard varieties with an ISO code. Thus, neither South Tyrolean nor Bavarian appear in any of the relevant multilingual evaluation benchmarks, such as translation-oriented benchmarks like FLORES-200 and NTREX-128, in the cross-lingual topic classification dataset SIB-200 [58], or in GlobalPIQA [59] or Global MMLU. The only benchmark that Bavarian is included in is the recent DialectBench [40] mentioned in Section 2.
Nevertheless, there has been some demand and language technology development in relation to South Tyrolean varities. More specifically, models are currently being fine-tuned specifically for automated transcription and subtitling of audio(visual) material [60]. The use case being addressed is that of transposing non-standard audio into standard German writing, implying that the aim is to allow broader access to non-standard audio(visuals) to speakers of Standard German and/or to provide a basis for machine translating South Tyrolean German into other languages. This, in turn, can be expected to have ramifications for speakers of South Tyrolean German. For instance, this might mean that they might be able to use non-standard forms in digitally mediated interaction with non-speakers of South Tyrolean German, in recorded interactions intended for a broader audience, such as TV broadcasts, or to voice-controlled digital assistants. However, as [19] has shown, whether such capabilities actually correspond to user needs and will be taken up in user practices is far from clear, as language technologies are ideologically associated with standard language varieties.
Kurdish is an Indo-European language belonging to the Northwestern Iranian branch, spoken by over 40 million people across Western Asia, primarily in Iraq, Turkey, Iran, Syria, and Armenia, as well as among substantial diaspora communities worldwide [61]. Rather than constituting a single unified language, Kurdish represents a dialect continuum comprising several distinct varieties with varying degrees of mutual intelligibility. The principal varieties include Northern Kurdish (Kurmanji), spoken by an estimated 15–20 million speakers predominantly in Turkey, Syria, northern Iraq, and northwestern Iran; Central Kurdish (Sorani), with approximately 10 million speakers concentrated in Iraqi Kurdistan and Iranian Kurdistan; and Southern Kurdish, spoken in Kermanshah, Ilam, and Lorestan [62]. Additionally, the Zaza-Gorani languages, including Zazaki and Hawrami, are spoken by communities who identify as ethnic Kurds, though their linguistic classification remains debated; some scholars group them within the broader Kurdish language family while others consider them closely related but distinct Northwestern Iranian languages [63].
| Variety | Supported Domains | Score (/4) |
|---|---|---|
| Central Kurdish (Sorani, ckb) | Education, Media, Press, Official | 4 |
| Northern Kurdish (Kurmanji, kmr) | Education, Media, Press, Official | 4 |
| Southern Kurdish (sdh) | Press, Publishing | 2 |
| Zazaki (zza) | Media, Publishing | 2 |
| Hawrami (Gorani, hac) | Publishing | 1 |
| Laki (lki) | None | 0 |
Divided across multiple nation-states, Kurdish has historically faced systematic suppression and assimilation campaigns, including outright bans on public use in Turkey (1923–1991) [64] and Persianization and Arabization policies targeting its use in formal contexts [65]. [66] documents how such policies have led Kurdish-speaking parents in certain regions to use the dominant state language with their children, contributing to intergenerational language shift. Nevertheless, the communicative spaces for Kurdish in media and education have expanded over the past decades, mainly thanks to the establishment of the Kurdistan Regional Government in Iraq, where Kurdish now enjoys official status and institutional support [67]. Unlike Northern Kurdish and Central Kurdish, other varieties such as Southern Kurdish, Hawrami, and Zazaki remain severely under-represented, with Hawrami classified as “definitely endangered” by UNESCO [68].
The orthographic landscape of Kurdish reflects its political fragmentation. Northern Kurdish is written using a Latin-based alphabet, while Central Kurdish employs an Arabic-based script. Soviet-era Kurdish communities used a Cyrillic-based system. This orthographic diversity, while enabling written expression within political boundaries, creates significant barriers to cross-dialectal communication, pan-Kurdish linguistic unity [69] and efficient language technology [70].
Table 2 summarizes the distribution of institutional support across Kurdish varieties in four key sociolinguistic domains: education, media, publishing, and official use. Central Kurdish and Northern Kurdish, the two most widely spoken varieties that are also standardised to some degree [62], exhibit full domain coverage, benefiting from formal recognition in educational curricula, broadcast media, print press, and governmental functions. In contrast, Southern Kurdish and Zazaki maintain a more limited presence, primarily confined to press and publishing activities, reflecting their exclusion from official and educational institutions. Hawrami receives support only in publishing despite ongoing efforts for its official recognition along with Central Kurdish and Northern Kurdish [71], [72]. As the least supported variety, Laki lacks institutional backing entirely. This uneven distribution underscores the hierarchical nature of language vitality within the Kurdish continuum, where political recognition and demographic weight correlate strongly with institutional investment, a disparity that poses significant challenges for language preservation efforts and the development of inclusive Kurdish language technologies.
| Variety | Data | Models | Machine Translation | |||||
| Bitext | Audio | XLM-R | BERT | MADLAD-400 | TranslateGemma | Microsoft | ||
| Central Kurdish | \(<\)300M | \(~\)200h | ✔ | ✔ | ✔ | ✔ | ||
| Northern Kurdish | \(<\)300M | \(~\)200h | ✔ | ✔ | ✔ | ✔ | ||
| Southern Kurdish | \(<\)10M | \(~\)10h | ||||||
| Laki | \(<\)1M | \(~\)2h | ||||||
| Zazaki | \(<\)10M | \(~\)2h | ||||||
| Hawrami | \(<\)10M | \(~\)20h | ||||||
| Variety | FLORES-200 | NTREX-128 | SIB-200 | GlobalPIQA | Global MMLU |
|---|---|---|---|---|---|
| Central Kurdish | ✔ | ✔ | ✔ | ✔ | |
| Northern Kurdish | ✔ | ✔ | ✔ | ||
| Southern Kurdish | |||||
| Laki | |||||
| Zazaki | |||||
| Hawrami |
The sociolinguistic situation of Kurdish has direct implications for its computational processing and representation in language and speech technologies. The historical suppression of the language resulted in limited written corpora, leading to Kurdish being consistently classified as a low-resource language in NLP research. A survey of the NLP literature carried out by AUTHOR reveals that among over 100 papers published in this field, only a handful address varieties other than Central and Northern Kurdish. To further illustrate this disparity, we assess resourcefulness from the three essential pillars of modern language technology: (1) data (text and audio) needed for training models, (2) pretrained models, including embeddings essential for representation, and (3) evaluation benchmarks. Additionally, given its importance as an NLP application, we consider machine translation support as an indicator of technological investment.
Table 3 presents an overview of available resources across Kurdish varieties in terms of parallel text corpora, audio data, pretrained language models, and machine translation services. Central Kurdish followed by
Northern Kurdish emerges as the most resourced variety, with over 300 million tokens of textual data and approximately 200 hours of transcribed audio, alongside support from multilingual models such as MADLAD-400 [73] and TranslateGemma [74], as well as commercial translation services from Google and Microsoft. Critically, Southern Kurdish, Laki, Zazaki, and Hawrami remain entirely unsupported across
all categories, reflecting a near-total absence from the modern NLP ecosystem.
Table 4 provides a complementary perspective by surveying the inclusion of Kurdish varieties in prominent multilingual evaluation benchmarks. Benchmarks are essential for assessing LLM performance in an era of rapid technological advancement. Central and Northern Kurdish appear in translation-oriented benchmarks such as FLORES-200 and NTREX-128, as well as in the cross-lingual topic classification dataset SIB-200 [58]. Central Kurdish is additionally represented in GlobalPIQA [59] for commonsense reasoning. However, neither variety is included in Global MMLU, and the remaining four varieties—Southern Kurdish, Laki, Zazaki, and Hawrami—are absent from all surveyed benchmarks entirely.
It should be noted that the models and benchmarks presented here are not intended to be exhaustive; rather, they are selected to highlight the stark discrepancy in computational support among varieties of Kurdish. This technological marginalization mirrors and reinforces the sociolinguistic hierarchies discussed earlier, posing substantial barriers to the development of inclusive, pan-Kurdish language technologies.
The persistent underrepresentation of Kurdish varieties in NLP can be attributed, in part, to a structural lack of investment from governments and policymakers. Unlike languages that benefit from state-sponsored digitization initiatives, corpus development programs, or dedicated research funding, Kurdish, particularly its lesser-resourced varieties, has received negligible institutional support for computational linguistics research. This absence of top-down investment stands in stark contrast to the resources allocated to dominant state languages in the regions where Kurdish is spoken [75].
In the absence of such funding, the development of modern LLMs has come to rely predominantly on publicly available web data, operationalizing the concept of the “Web as corpus” [76]. However, this approach inherently disadvantages languages with limited digital presence, creating a feedback loop in which low online visibility leads to exclusion from training corpora, which in turn perpetuates technological marginalization. For Kurdish, the historical suppression of the language in public and institutional domains has directly constrained its digital footprint, leaving vast gaps in web-crawled datasets.
Moreover, where data collection efforts do exist, whether by individual researchers, diaspora communities, or private enterprises, the resulting resources frequently remain proprietary or unpublished. This reluctance or inability to release data under open-source licenses compounds the scarcity problem, as subsequent research cannot build upon prior work. The result is a vicious circle: the lack of open resources discourages new investment, while the absence of investment limits the creation of shareable datasets. Until this cycle is disrupted through coordinated open-source initiatives, sustained funding, or policy intervention, the majority of Kurdish varieties will remain stranded at the margins of language technology development.
The development of LLMs for non-standard languages like South Tyrolean German represents more than a technical challenge; it is a battleground for linguistic and digital sovereignty. As we have argued, the digital language divide is maintained by a complex interplay of market forces, historical state policies, and the inherent biases of standardized data pipelines. The cases of South Tyrolean and Kurdish varieties illustrate how linguistic hierarchies, rooted in standardization, colonialism, and scriptism, are reproduced and intensified through tokenization practices, morphological biases, and evaluation frameworks that privilege standardised, high‑resource languages. Ultimately, closing the dialect gap requires a multi-faceted approach. We advocate for:
The implementation of dialect gap reporting and Corporate Social Responsibility (CSR) credits to incentivize support for under-resourced varieties.
Adopting the principle of “nothing about us without us,” ensuring that communities retain data sovereignty and lead the creation of modern technological terminology.
Leveraging the expertise of academics and language institutes to provide the sociolinguistic granularity that prevents “hallucinated standardization.”
By moving away from extractive data practices and toward participatory stewardship, we can ensure that AI serves to revitalize rather than erase non-standard linguistic repertoires.
Similar results hold for tokenizer in more recent LLMs. OpenAI provide a platform for this at https://platform.openai.com/tokenizer↩︎
[1] reports that most languages pay \(2-3\times\) the price of English, with languages that use non-Latin scripts such as Arabic paying even higher costs up to \(5\times\) and severely under-resourced languages such as Odia paying over \(15\times\) the cost↩︎
Paraphrasing the famous quote by “what gets measured gets managed” by Peter Drucker↩︎