Comparative Performance Analysis of Xception and ResNet50 Architectures for Facial Expression Recognition Using the FER-2013 Dataset
Diah Aryani(1*), Habibullah Akbar(2), Ferdinand Defin Delio(3)
(1) Universitas Esa Unggul
(2) Universitas Esa Unggul
(3) Universitas Esa Unggul
(*) Corresponding Author
Abstract
This study aims to evaluate and compare the performance and computational efficiency of the Xception and ResNet50 architectures in facial expression classification tasks. Facial Expression Recognition (FER) plays an important role in the development of intelligent systems capable of interpreting human emotions. This technology has various applications, including adaptive learning systems, emotion-aware customer service, and mental health support systems. This research compares two widely used Convolutional Neural Network (CNN) architectures, Xception and ResNet50, for facial expression classification using the FER-2013 dataset. The dataset contains 35,887 grayscale facial images with a resolution of 48×48 pixels categorized into seven basic emotions. All images were resized to 224×224 pixels and converted into RGB format to match the input requirements of pretrained ImageNet models.
Both architectures were trained using a transfer learning strategy with selective fine-tuning on specific layers. Data augmentation techniques were applied to increase dataset variability and reduce overfitting. Model performance was evaluated using accuracy, precision, recall, F1-score, and confusion matrix metrics. The results show that Xception architecture outperforms ResNet50, achieving a validation accuracy of 70.69% and a weighted F1-score of 0.71. These findings demonstrate that appropriate architecture selection and structured training strategies can significantly improve FER performance in practical intelligent systems.
Keywords
Full Text:
PDFReferences
REFERENCES
F. Chollet, “Xception: Deep Learning with Depthwise Separable Convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 1800–1807. doi: 10.1109/CVPR.2017.195.
R. K. Kumar and S. Vuppala, “Robust Facial Expression Recognition using CNN-based Feature Fusion,” IEEE Access, vol. 9, pp. 49385–49395, 2021, doi: 10.1109/ACCESS.2021.3069164.
J. He, “Face Expression Recognition (FER) Dataset.” 2019. [Online]. Available: https://www.kaggle.com/datasets/jonathanoheix/face-expression-recognition-dataset
Y. Gao, Y. Cai, X. Bi, B. Li, S. Li, and W. Zheng, “Cross-Domain Facial Expression Recognition through Reliable Global–Local Representation Learning and Dynamic Label Weighting,” Electronics, vol. 12, no. 21, 2023, doi: 10.3390/electronics12214553.
Z. Zhao and X. Zhang, “Static Facial Expression Recognition via Multi-scale CNN,” J. Vis. Commun. Image Represent., vol. 76, 2021, doi: 10.1016/j.jvcir.2021.103115.
A. Dapogny, K. Bailly, and S. Adam, “Confidence-based Adversarial Training for Robust Facial Expression Recognition,” IEEE Trans. Affect. Comput., 2021.
S. M. J. S. Priya and S. Umamaheswari, “Deep Convolutional Neural Network-Based Facial Expression Recognition,” Expert Syst. Appl., vol. 197, 2022, doi: 10.1016/j.eswa.2022.116696.
H. Sheng and M. Lau, “Optimising Facial Expression Recognition: Comparing ResNet Architectures for Enhanced Performance,” in International Conference on Control, Dynamic Systems and Robotics, 2024. doi: 10.11159/cdsr24.123.
M. Kaur, “Facial Emotion Recognition: A Comprehensive Review,” Expert Syst., vol. 41, no. 10, pp. 105–112, 2024, doi: 10.1111/exsy.13670.
A. Mollahosseini, B. Hasani, and M. H. Mahoor, “AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild,” IEEE Trans. Affect. Comput., vol. 10, no. 1, pp. 18–31, 2019, doi: 10.1109/TAFFC.2017.2740923.
J. Yu, Y. Liu, and R. Fan, “MixCut: A Data Augmentation Method for Facial Expression Recognition.” 2024. [Online]. Available: http://arxiv.org/abs/2405.10489
S. L. Li and W. Deng, “Deep Facial Expression Recognition: A Survey,” IEEE Trans. Affect. Comput., vol. 13, no. 3, pp. 1195–1215, 2022, doi: 10.1109/TAFFC.2020.2981446.
W. Li and J. He, “Cross-dataset Facial Expression Recognition via CNN Feature Alignment,” Pattern Recognit. Lett., vol. 125, pp. 166–171, 2019, doi: 10.1016/j.patrec.2019.01.009.
A. Shanimol and J. Charles, “ResNet50 and GRU: A Synergistic Model for Accurate Facial Emotion Recognition,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 8, pp. 611–620, 2024, doi: 10.14569/IJACSA.2024.0150861.
L. Liao, S. Wu, C. Song, and J. Fu, “RS-Xception: A Lightweight Network for Facial Expression Recognition,” Electronics, vol. 13, no. 16, 2024, doi: 10.3390/electronics13163217.
J. Yang, L. Tang, Z. Chen, X. Zhu, and X. Lai, “Facial Emotion Recognition Based on Optimized Xception Training,” in ACM International Conference Proceeding Series, 2024. doi: 10.1145/3697355.3697358.
C. W. Wee and S. Budianto, “Improving Facial Expression Recognition with Attention-guided CNN,” Image Vis. Comput., vol. 103, 2020.
Y. Bengio, A. Courville, and P. Vincent, “Representation Learning: A Review and New Perspectives,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 35, no. 8, pp. 1798–1828, 2013.
Article Metrics
Copyright (c) 2026 IJCCS (Indonesian Journal of Computing and Cybernetics Systems)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
View My Stats1






