AI & Computingarticle2026-08-28

Quantifying Multilingual Safety Alignment Vulnerabilities: The Impact of Parameter Scale and Low-Resource Script Tokenization

Open access0 citations

Abstract

Safety alignment in Large Language Models relies heavily on Reinforcement Learning from Human Feedback and system instruction formatting. However, existing safety datasets and automated red-teaming benchmarks remain predominantly English-centric, creating severe evaluation blind spots when models process regional low-resource languages and non-Latin scripts. This paper presents an empirical analysis measuring the safety resilience of open-weights models across parameter scales, specifically comparing GPT-OSS-20B and GPT-OSS-120B against a standardized testbed of 50 security vulnerability prompt vectors. Each prompt was evaluated across three linguistic variants: English, native Urdu in Nastaliq script, and native Punjabi in Shahmukhi and Gurmukhi scripts, yielding 300 trial iterations. Our findings demonstrate an inverse relationship between parameter scaling and jailbreak compliance: the 20B model yielded an overall compliance rate of 81.33%, whereas the heavily aligned 120B model decreased to 64.00% (χ² = 10.4889, p = 0.0012). Furthermore, we document a verbosity-scale paradox where larger models provide significantly denser technical detail upon bypass, alongside cross-lingual alignment divergence in 30% of tested configurations due to script tokenization fragmentation.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-28

Authors: Faizan e Bahoo Chaudhry

Institutions: Government College University, Lahore