Researchers tested three popular AI language systems in 12 languages from five language families: Indo-European, Afro-Asiatic, Turkic, Sino-Tibetan and Japonic. The systems showed substantial accuracy across languages with different structures, but they did not match human performance equally closely in every language.

The results challenged the expectation that English is always the strongest language for these systems. Several Romance languages outperformed English, including languages the researchers describe as lower-resource. The authors discuss tokenization, the amount and origin of training data, and how closely a language is related to Spanish and English as possible influences on performance.