Privacy-preserving federated learning for detection of AI-generated text using hybrid RoBERTa-CNN models in low-resource languages
Abstract
The rapid adoption of large language models (LLMs) has transformed educational writing while introducing significant challenges for academic integrity and authorship verification. Existing AI-generated text detection systems predominantly rely on centralized cloud-based infrastructures, requiring student essays to be uploaded to external servers and raising concerns regarding privacy, data ownership, and regulatory compliance. To address these challenges, we propose RoBERTa-CNN+FL, a privacy-preserving framework that integrates a RoBERTa encoder for contextual semantic representation, a convolutional neural network (CNN) for local stylistic feature extraction, federated learning (FL) for decentralized collaborative model training, and differential privacy (DP) to protect sensitive educational data. The framework was evaluated on the LowResourceQAEssaysText dataset comprising human-written and AI-generated essays generated by multiple contemporary LLMs. Performance was assessed using accuracy, precision, recall, F1-score, Matthews correlation coefficient, ROC-AUC, BLEU score, perplexity, and human evaluation conducted by independent academic reviewers to assess the consistency of model predictions with expert judgments. Experimental results demonstrate that the proposed hybrid framework achieves robust and reliable detection performance while preserving data privacy and enabling on-device deployment. Comparative experiments further show competitive performance against widely used commercial and research AI-text detection systems, including Turnitin, Grammarly AI Detector, QuillBot, GPTZero, Copyleaks, Originality.ai, Sapling AI Detector, Writer.com Detector, DetectGPT, Fast-DetectGPT, Crossplag AI Detector, ZeroGPT, OpenAI AI Classifier, Hugging Face open-source detectors, and other academic detection frameworks. The proposed framework offers a scalable and privacy-preserving solution for AI-generated text detection in multilingual and low-resource educational environments, reducing dependence on centralized cloud-based services while supporting trustworthy academic integrity assessment.
// Source
Authors: Manish Prajapati, Santos Kumar Baliarsingh, Ramasubbareddy Somula
Institutions: Symbiosis International University, KIIT University, Visvesvaraya Technological University