A study of three popular systems across 12 languages found that English was not consistently their strongest language.
Researchers tested three popular AI language systems in 12 languages from five language families: Indo-European, Afro-Asiatic, Turkic, Sino-Tibetan and Japonic. The systems showed substantial accuracy across languages with different structures, but they did not match human performance equally closely in every language.
The results challenged the expectation that English is always the strongest language for these systems. Several Romance languages outperformed English, including languages the researchers describe as lower-resource. The authors discuss tokenization, the amount and origin of training data, and how closely a language is related to Spanish and English as possible influences on performance.
How the languages compared
The researchers prompted three popular AI language systems in 12 languages spanning five language families. The systems showed high accuracy across languages with widely differing structures, but their performance varied and could fall below human baselines by different amounts.
English was not consistently the best-performing language. Several Romance languages systematically outperformed it, including some described as lower-resource. The researchers identify tokenization, language distance from Spanish and English, training-data size, and whether data came from high- or low-resource and WEIRD or non-WEIRD communities as factors that may help explain the differences.
Why language coverage matters
Many evaluations of AI language systems focus on high-resource languages, especially English. These results indicate that English performance cannot automatically be used as a guide to how well a system will understand another language, including some languages with fewer digital resources.
That matters for people using AI systems to access information in different languages. Evaluating a wider range of languages can reveal uneven performance that English-focused tests may miss.
Evidence and limits
The study is a comparative evaluation based on prompts given to three popular AI language systems in 12 languages. It compares the systems' accuracy with human baselines and examines differences across language families and resource levels.
The abstract does not report the specific tasks, scores or human-baseline results. The study covers only three systems and 12 languages, so its findings should not be assumed to apply to every AI language system or every natural language. The proposed factors behind performance differences are presented as influences to consider, rather than as effects established by the comparison alone.
// Source
Scientific Reports · 2026 · DOI: 10.1038/s41598-026-69546-8
Authors: Natalia Moskvina, Raquel Montero, Masaya Yoshida, Ferdy Hubers, Paolo Morosi, Walid Irhaymi, Yan Jin, Tamara Serrano, Elena Pagliarini, Fritz Günther, Evelina Leivada
Institutions: Universitat Autònoma de Barcelona, Radboud University Nijmegen, Institució Catalana de Recerca i Estudis Avançats, Humboldt-Universität zu Berlin