AI & Computingarticle2026-08-15

MULTI-FUSE: Multimodal Late Fusion Framework for Hate Speech Detection in Videos

Open access0 citations

Abstract

Abstract Addressing cyberbullying and fostering more welcoming online communities depend on the detection of hateful content. However, because it requires examin-ing a number of interrelated components, including spoken words, visual imagery, and text within those visuals, detecting hate speech in videos is difficult. Thiswork presents MULTI-FUSE, a state-of-the-art hate speech detection model thatsuccessfully combines multimodal features from video content using a late fusion technique. The audio and video frames are processed separately by the frame-work. It analyzes frames for hate symbols, text (using OCR), and gestures to generate a visual hate score using deep learning techniques, while extractingand analyzing speech to generate an audio hate score. These audio and visualhate scores are then combined into a single feature vector, which is input into asupervised learning classifier for the final prediction of hate speech. This studypresents an optimized multimodal framework for detecting hate speech in videos by assessing various supervised learning models and identifying the Support Vec-tor Machine (SVM) as the most effective classifier, achieving an accuracy of 0.8571, precision of 0.8333, recall of 0.8929, and F1-score of 0.8621. The frame-work systematically determines the optimal frame-processing rate of one frame per second to strike a balance between accuracy and computational efficiency. Anablation study confirms that the full multimodal combination (audio, text, andvisual) significantly outperforms all unimodal configurations, achieving a 7.5%relative improvement over the best unimodal model. Compared with reportedbaselines on the HATEMM dataset, MULTI-FUSE achieves an 18% improvementin F1-score over existing late-fusion approaches. The late fusion method mergesaudio, text, and visual hate scores, resulting in strong detection performance

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-15

Authors: Lekshmi M S and Malathy S

Institutions: Karpagam Academy of Higher Education