ISSN : 3049-1754 (PRINT) ISSN : 3048-8966 (ONLINE)
DEEPFAKE DETECTION USING XCEPTION MODEL WITH MULTI-MODAL ENSEMBLE LEARNING AND BAYESIAN FUSION
Affiliation
Monash University, Caulfield, Australia
Abstract:
Deepfake technology, powered by generative adversarial networks (GANs) and autoencoders, has advanced to a stage where synthetic media is increasingly indistinguishable from authentic content, posing significant threats to digital media integrity, electoral processes, financial systems, and personal reputation. This paper presents a robust deepfake detection framework that combines multimodal ensemble learning with Bayesian fusion to address the limitations of unimodal detectors. The proposed system integrates an Xception-based image stream fine-tuned on FaceForensics++, a DeepFace emotion analysis stream that captures unnatural affective cues, and a SlowFast temporal stream that exploits spatio-temporal inconsistencies across video frames. The three streams are combined using a Bayesian fusion module implemented in Pyro, which assigns probabilistic weights to each modality and explicitly models predictive uncertainty. Experimental evaluation on FaceForensics++ using fivefold cross-validation shows that the proposed ensemble achieves 92.3% accuracy and an F1-score of 0.92, outperforming the Xception baseline by 5.1 percentage points and reducing the false positive rate from 9.2% to 6.5%. The system operates at 40 FPS on GPU and 8.3 FPS on CPU, demonstrating real-time feasibility. The study contributes a principled, uncertainty-aware fusion architecture for deepfake detection and discusses its practical implications for media forensics, journalism, social media moderation, and law enforcement.
Keywords:
Deepfake Detection, Ensemble Learning, Bayesian Fusion, Xception, SlowFast, DeepFace, Uncertainty Modelling, Multi-Modal Learning.
Publishing Chronology:
Received - 18/01/2026
Revised - 15/02/2026
Accepted - 21/04/2026
References:
Afchar, D., Nozick, V., Yamagishi, J., & Echizen, I. (2018, December). Mesonet: a compact facial video forgery detection network. In 2018 IEEE international workshop on information forensics and security (WIFS) (pp. 1-7). IEEE.DOI: 10.1109/WIFS.2018.8630761
Chollet, F. (2017). Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1251-1258).
Dolhansky, B., Howes, R., Pflaum, B., Baram, N., & Ferrer, C. C. (2020). The Deepfake Detection Challenge dataset. arXiv preprint arXiv:2006.07397.
Gal, Y., & Ghahramani, Z. (2016). Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. Proceedings of the 33rd International Conference on Machine Learning (ICML), 1050–1059.
Gal, Y., & Ghahramani, Z. (2016, June). Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning (pp. 1050-1059). PMLR.
Heo, Y., Choi, Y., Lee, Y., & Kim, H. (2023). Deepfake detection via combining frequency and spatial features with vision transformer. Sensors, 23(7), 3489.
Khan, S. A., & Dai, H. (2021, October). Video transformer for deepfake detection with incremental learning. In Proceedings of the 29th ACM international conference on multimedia (pp. 1821-1828).
Qi, H., Guo, Q., Juefei-Xu, F., Xie, X., Ma, L., Feng, W., ... & Zhao, J. (2020, October). Deeprhythm: Exposing deepfakes with attentional visual heartbeat rhythms. In Proceedings of the 28th ACM international conference on multimedia (pp. 4318 4327).
Liu, Y., Zhu, X., Zhao, X., & Cao, Y. (2023). Spatial-phase shallow learning for deepfake detection. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 45(3), 2793–2807.
Mirsky, Y., & Lee, W. (2021). The creation and detection of deepfakes: A survey. ACM Computing Surveys (CSUR), 54(1), 1–41.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., ... & Chintala, S. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
Sabir, E., Cheng, J., Jaiswal, A., AbdAlmageed, W., Masi, I., & Natarajan, P. (2019). Recurrent convolutional strategies for face manipulation detection in videos. Interfaces (GUI), 3(1), 80-87.
Singh, A. (2020, March 15). How generative adversarial networks (GAN) work. DataDrivenInvestor.
Tan, C., Zhao, Y., Wei, S., Gu, G., Liu, P., & Wei, Y. (2024, March). Frequency-aware deepfake detection: Improving generalizability through frequency space domain learning. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, No. 5, pp. 5052-5060).
Wightman, R. (2019). PyTorch image models (timm) [Computer software]. GitHub repository.
Yan, Z., Zhang, Y., Yuan, X., Lyu, S., & Wu, B. (2023). Deepfakebench: A comprehensive benchmark of deepfake detection. arXiv preprint arXiv:2307.01426.
Zhao, T., Xu, X., Xu, M., Ding, H., Xiong, Y., & Xia, W. (2021). Learning self consistency for deepfake detection. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 15023-15033).
We use cookies to analyze website traffic and optimize your website experience. By accepting our use of cookies, your data will be aggregated with all other user data.