TTLD: transformer-based threatening language detection in urdu social media text
Abstract
Organizations now treat social media activity as a live indicator of public sentiment, useful for decisions that depend on knowing what is happening in real time. The same platforms, however, also carry threatening language at scale, and that has direct consequences for public safety. Urdu sits at a disadvantage here: it is morphologically complex, resource-poor in NLP terms, and has received little attention in automated threat-detection work to date. We first benchmark pre-trained language models spanning multilingual, monolingual, and Twitter-specific variants on this task, then use what that comparison reveals to design TTLD (Transformer-based Threatening Language Detection). TTLD combines XLM-T for domain-adapted contextual embeddings, a Bi-LSTM component for sequential dependencies, and multi-head attention for weighting informative input segments. On an annotated Urdu tweet dataset, TTLD reaches an F1-score of 90.42%, ahead of every baseline considered here, suggesting a practical path toward better-supported threat detection in Urdu and other low-resource social media settings.
// Source
Authors: Wahab Khan, Rafiul Haq, Ihsanullah Khan, Aurangzeb Khan, Muhammad Mansoor Alam, Mazliham Bin Mohd Su’ud
Institutions: Riphah International University, Multimedia University, Tianjin University, University of Science and Technology Bannu