Society & Economicspreprint2026-08-15

Crossing the Sampler Boundary: When a Foreign-Script Token Becomes Decodable

Open access0 citations

Abstract

We study a narrow failure mode in multilingual language models: isolated foreign-script tokens appearing inside otherwise monolingual generation. We separate five stages that are often conflated: local token margin, vocabulary rank, sampler eligibility, post-warp probability, and actual sampling. On 98,106 English next-token states from Qwen3-0.6B, we reconstruct the exact temperature/top-k/top-p decoding transform and show that rank-based metrics can move in the opposite direction from sampler eligibility under the same inference-time perturbation. A Q4g32 perturbation reduces Han top-20 membership while increasing exact sampler eligibility under several decoding configurations. Some newly eligible foreign tokens are semantically coherent cross-language competitors of the English continuation. Monte Carlo sampling and live generation validate the predicted post-warp probabilities. The effect replicates on Qwen2.5-0.5B, is absent on BLOOM-560m, appears across several writing systems, and varies non-monotonically across quantization schemes and perturbation strengths. The results support a narrow claim: small local cross-language boundaries can be crossed in either direction by ordinary inference-time perturbations, and decoding configuration is a major determinant of whether those crossings become observable outputs.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-15

Authors: Arseniy Abramidze

Institutions: Peter the Great St. Petersburg Polytechnic University