AI & Computingarticle2026-08-05

Trustworthy Foundation Model-Based Multimodal Biometric Authentication for Secure and Privacy-Preserving Intelligent Systems Using Face, Ear Lobe, and Iris

Open access0 citations

Abstract

The increasing deployment of intelligent systems in banking, healthcare and critical infrastructure emphasizes the need for authentication procedures that are accurate, spoofing-robust, and privacy-preserving. In this paper, we propose a Trustworthy Foundation Model based Multimodal Biometric Authentication (TFM-MBA) framework to fuse three complementary biometric modalities, i.e., face, ear lobe and iris, via representations extracted from a large pre-trained vision foundation model, which are adapted to each modality with lightweight modality-specific adapters. Specifically, our architecture is built around: an attention-based fusion module; a trust-calibration layer for estimating the confidence and modality reliability of each sample; and a privacy-preserving pipeline combining federated learning, differential privacy and partial homomorphic encryption for template protection. We evaluate the system on combined benchmark-style datasets (face: CASIA-WebFace/LFW-style splits; ear: AWE/IITD-Ear-style splits; iris: CASIA-Iris-style splits) with a unified multimodal protocol with simulated real-world degradations (occlusion, blur, illumination change and presentation attacks). The experimental results demonstrate that the proposed multimodal fusion achieves a rank-1 accuracy of 99.1% and an Equal Error Rate (EER) of 0.18%, exceeding the best unimodal baseline (iris-only, 97.4% accuracy, 0.71% EER) and current multimodal fusion baselines by 1.2-3.6 percentage points in accuracy. The privacy-preserving version of the pipeline keeps 98.6% of the non-private model’s accuracy while reducing the membership-inference attack success rate from 68.3% to 52.1%, close to random guessing. We further show that the trust-calibration module improves the spoof-detection F1-score to 0.978 and offers interpretable per-modality contribution scores, which improve system transparency. The results demonstrate that the integration of foundation-model representations with modality-aware fusion and privacy-preserving training results in a biometric authentication system that is accurate, robust, explainable, and in line with data-protection standards for real-world intelligent systems.

// Source

View paper (DOI)Open access versionOpenAlexZenodo (CERN European Organization for Nuclear Research)Published 2026-08-05

Authors: ARUNARANI SHANMUGAMANI

Institutions: SRM University