Quantifying Multilingual Safety Alignment Vulnerabilities: The Impact of Parameter Scale and Low-Resource Script Tokenization
Abstract
Safety alignment in Large Language Models relies heavily on Reinforcement Learning from Human Feedback and system instruction formatting. However, existing safety datasets and automated red-teaming benchmarks remain predominantly English-centric, creating severe evaluation blind spots when models process regional low-resource languages and non-Latin scripts. This paper presents an empirical analysis measuring the safety resilience of open-weights models across parameter scales, specifically comparing GPT-OSS-20B and GPT-OSS-120B against a standardized testbed of 50 security vulnerability prompt vectors. Each prompt was evaluated across three linguistic variants: English, native Urdu in Nastaliq script, and native Punjabi in Shahmukhi and Gurmukhi scripts, yielding 300 trial iterations. Our findings demonstrate an inverse relationship between parameter scaling and jailbreak compliance: the 20B model yielded an overall compliance rate of 81.33%, whereas the heavily aligned 120B model decreased to 64.00% (χ² = 10.4889, p = 0.0012). Furthermore, we document a verbosity-scale paradox where larger models provide significantly denser technical detail upon bypass, alongside cross-lingual alignment divergence in 30% of tested configurations due to script tokenization fragmentation.
// Source
Authors: Faizan e Bahoo Chaudhry
Institutions: Government College University, Lahore