Asking an AI model the same question in Arabic, English or French can produce answers that vary in the accuracy of their information, terminology and detail, as well as in their understanding of cultural context. This does not necessarily mean the model “thinks” better in one language than another. It is mainly linked to how it was trained, the volume of text available in each language and how languages are represented computationally within its architecture.
Differences in training data between languages
The differences begin with the data used to train large language models. Although vast quantities of text are used, these materials are not distributed equally among languages.
The study linked these differences to the volume and sources of training data, the linguistic distance between languages and how texts are split before processing. The results also showed that English did not lead in every test, with other languages outperforming it on some tasks.
Another study covering more than 250 languages found that incorporating multilingual data into training can improve the performance of low-resource languages. However, the extent of the improvement depends on the volume of available data and the degree of linguistic similarity between languages.
How word tokenisation affects model performance
Conversely, expanding excessively the number of languages handled by a model may constrain its capabilities rather than produce equal improvements across all languages. The issue is not limited to the volume of text, because words do not enter the model in their usual form. Before a question is processed, texts are divided into smaller units known as tokens, and the efficiency of this process may vary from one language to another.
As a result, two questions identical in meaning can reach the model in two different computational forms when written in different languages. Recent research has indicated that the choice of tokeniser affects both model performance and training costs. Tools designed around English may cause a significant decline when used with other languages because they represent texts in those languages less efficiently.
This technical shortcoming translates into differences in understanding questions, formulating answers and retrieving information. A study presented at the 2025 conference of the Association for Computational Linguistics also concluded that linguistic diversity affects how text is tokenised, and that the resulting differences may affect model performance across multiple tasks.
Cultural context changes how a question is interpreted
Changing the language of a question is therefore not merely a matter of replacing one set of words with another. It may alter the number of units processed by the model and the computational relationships between them. Language also carries an associated culture, and some questions contain cultural references even when their wording appears simple.
A study presented at ACL 2025 identified what the researchers called “linguistic-cultural synergy”, in which models performed better when a question was culturally compatible with the language used. The researchers concluded that model evaluation should not separate language from its cultural context. Another study showed that models may possess knowledge of local cultures but do not automatically draw on that knowledge every time they answer in a language associated with that culture.
Explicitly including cultural context in a question can make the answer more closely connected to the local environment than a general question that does not specify the country or community concerned. This effect can be seen, for example, in questions about a social custom, a proverb or a legal system. The language used may suggest a particular country or culture to the model, which then builds its answer around those implicit cues.
Challenges for Arabic and low-resource dialects literacy
When the language of a question changes, the cues on which the model relies may change with it, even if the question’s direct meaning remains the same. The problem is more pronounced in languages and dialects with limited digital resources. In Arabic, the challenges go beyond handling Modern Standard Arabic to include the wide diversity of dialects and local contexts.
This makes building a model capable of delivering a comparable level of accuracy and cultural awareness across different Arab environments more complex. According to the “Balsam Arabic Language AI Technology Maturity Index” paper, published in the proceedings of the 2025 Arabic Language Processing Conference, model performance in Arabic is affected by several factors, including data scarcity, linguistic and dialectal diversity and morphological complexity.
The paper added that limited measurement tools and specialised tests pose another challenge to accurately assessing these models’ capabilities. Another study developed the “Aradis” benchmark to test models’ abilities in Arabic dialects and their awareness of cultural contexts in the Gulf, Egypt and the Levant.
The benchmark aims to address the gap between Modern Standard Arabic and dialects, alongside cultural differences within Arab regions, rather than assessing Arabic performance as a single linguistic block. Other languages face similar challenges. UNESCO has noted that AI systems trained mainly on dominant languages may be less effective when dealing with some African languages.
The impact of a shortage of local data is not limited to the quality of linguistic formulation. It may also limit systems’ ability to understand and process local cultural concepts. A difference between answers in two languages, on its own, is not conclusive evidence of deliberate bias.
The variation may arise from differences in data volume, text-tokenisation methods, cultural context or the wording of the question. Nevertheless, these factors can become actual linguistic bias in outputs, particularly when a language is represented by less data or performance tests are built around the most widely used languages.
Research is therefore moving towards testing models across multiple languages and cultures rather than assessing their performance solely in English. Testing the ability to understand words alone is not enough, because the quality of an answer also depends on understanding the question within its linguistic and cultural environment.
Experts say that a model does not retrieve a fixed answer from a single database. Instead, it builds one from representations formed during training. Its route to an answer therefore changes whenever the language, context or available data changes.