DOI QR코드

DOI QR Code

Identification Method for Human and AI-Synthesized Speech Using a Multi-Feature Fusion-Based Deep Learning Model

다중특징 융합 기반 딥러닝 모델을 이용한 인간의 실제 음성과 AI 합성 음성의 식별 방법

  • HongJae Jang (Department of Computer Engineering, Gachon University) ;
  • SeongHoon Kim (Research Center, Neowine Co., Ltd.) ;
  • GiTae Han (Department of Computer Engineering, Gachon University)
  • 장홍재 (가천대학교 컴퓨터공학과) ;
  • 김성훈 ((주)네오와인 연구소) ;
  • 한기태 (가천대학교 컴퓨터공학과)
  • Received : 2025.01.14
  • Accepted : 2025.05.23
  • Published : 2025.08.05

Abstract

The rapid advancement of Artificial Intelligence(AI) has achieved remarkable progress in various fields such as image editing, audio generation, and video manipulation. However, it has also introduced new security threats, including deepfake speech and voice spoofing. This paper proposes a multi-feature deep learning based AI synthetic speech detection method capable of addressing these threats with high accuracy. The ASVspoof 2021 dataset's Logical Access(LA) and Deepfake(DF) data were used for training and testing. The proposed system utilizes two audio features, Mel-Spectrogram and MFCC(Mel-Frequency Cepstral Coefficients), to convert audio data into visual and sequential forms for training and inference. To demonstrate the superiority of the proposed method, a comparative analysis was conducted with various models such as CNN, BiLSTM, Transformer, and ensemble methods. Experimental results showed that the multi-feature fusion model outperformed single and ensemble models. The proposed multi-feature fusion model, which combines ConvNeXt-base and BiLSTM using a Late Fusion approach, achieved the highest performance with an accuracy of 98.44 %. The method proposed in this paper is expected to serve as a key technology in future AI deepfake synthetic speech detection systems.

Keywords

Acknowledgement

이 논문은 2025년도 정부(과학기술정보통신부)의 재원으로 정보통신기획평가원의 지원을 받아 수행된 연구임(RS-2022-II220050, 데이터 플로우 구조 기반 PIM의 실행 및 프로그래밍 모델 개발).

References

  1. Seo, J., Current Status, Type, Trend, and Response Implications of Voice Phishing, Statistics Korea, 2022. https://kostat.go.kr/synap/skin/doc.html?fn=2f1b5a46a70ffa293ac2113cbd1b9635e28497f762079ae221b0d41b4129ca9d&rs=/synap/preview/board/12312/.
  2. Bekmanova, G.; Yelibayeva, G.; Yergesh, B.; Orynbay, L.; Sairanbekova, A.; Kaderkeyeva, Z., "Emotional Coloring of Kazakh People's Names in the Semantic Knowledge Database of 'Fascinating Onomastics' Mobile Application," Proceedings of International Conference, pp. 666-671, 2022.
  3. Jothimani, S.; Premalatha, K., "MFF-SAug: Multi Feature Fusion with Spectrogram Augmentation of Speech Emotion Recognition Using Convolution Neural Network," Chaos, Solitons & Fractals, p. 112512, 2022.
  4. Kim, S. C.; Lee, J. W.; Cho, K. O.; Park, J. G.; Oh, Y. T., "A Study on the Algorithm for Speech Recognition," The 39th Summer Conference of the Korean Institute of Electrical Engineers, pp. 2255-2256, 2008.
  5. Salur, M. U.; Aydın, İ., "A Soft Voting Ensemble Learning-Based Approach for multimodal Sentiment Analysis," Neural Computing & Applications, pp. 18391-18406, 2022. https://doi.org/10.1007/s00521-022-07451-7
  6. Jeon, B.-U.; Kang, J.-S.; Chung, K., "AutoML and CNN-Based Soft-Voting Ensemble Classification Model for Road Traffic Emerging Risk Detection," Journal of Convergence Information Technology, pp. 14-20, 2021.
  7. J. Gondohanindijo, Muljono, E. Noersasongko, Pujiono, and D. R. M. Setiadi, "Multi-Features Audio Extraction for Speech Emotion Recognition Based on Deep Learning," International Journal of Advanced Computer Science and Applications, Vol. 14, No. 6, 2023, doi: https://doi.org/10.14569/ijacsa.2023.0140623.
  8. C. Wang, Y. Ren, N. Zhang, F. Cui, and S. Luo, "Speech emotion recognition based on multi‐feature and multi‐lingual fusion," Multimedia Tools and Applications, Aug. 2021, doi: https://doi.org/10.1007/s11042-021-10553-4.
  9. T. Baltrusaitis, C. Ahuja, and L.-P. Morency, "Multimodal Machine Learning: A Survey and Taxonomy," IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 41, No. 2, pp. 423-443, Feb. 2019, doi: https://doi.org/10.1109/tpami.2018. 2798607.
  10. S. Y. Boulahia, A. Amamra, M. R. Madi, and S. Daikh, "Early, intermediate and late fusion strategies for robust deep learning-based multimodal action recognition," Machine Vision and Applications, Vol. 32, No. 6, Sep. 2021, doi: https://doi.org/10.1007/s00138-021-01249-8.
  11. B.-U. Jeon, J.-S. Kang, and K. Chung, "AutoML and CNN-based Soft-voting Ensemble Classification Model For Road Traffic Emerging Risk Detection," Journal of Convergence Information Technology, Vol. 11, No. 7, pp. 14-20, Jan. 2021, doi: https://doi.org/10.22156/cs4smb.2021.11.07.014.
  12. Liu, Z.; Mao, H.; Wu, C.-Y.; Feichtenhofer, C.; Darrell, T.; Xie, S., "A ConvNet for the 2020s," Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR), 2022.
  13. A. Dosovitskiy et al., "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale," in International Conference on Learning Representations, 2021.
  14. Schuster, M.; Paliwal, K. K., "Bidirectional Recurrent Neural Networks," IEEE Transactions on Signal Processing, pp. 2673-2681, 1997.
  15. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; Polosukhin, I., "Attention Is All You Need," 2017.
  16. Benesty, J.; Sondhi, M. M.; Huang, Y., Eds., Springer Handbook of Speech Processing, Springer, Berlin, 2008.
  17. Allen, J., "Short Term Spectral Analysis, Synthesis, and Modification by Discrete Fourier Transform," IEEE Transactions on Acoustics, Speech, and Signal Processing, pp. 235-238, 1977.
  18. Davis, S.; Mermelstein, P., "Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences," IEEE Transactions on Acoustics, Speech, and Signal Processing, pp. 357-366, 1980.
  19. Logan, B., "Mel Frequency Cepstral Coefficients for Music Modeling," Proceedings of the 1st International Symposium on Music Information Retrieval, 2000.
  20. LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P., "Gradient-Based Learning Applied to Document Recognition," Proceedings of the IEEE, pp. 2278-2324, 1998.
  21. Delgado, H.; Evans, N.; Kinnunen, T.; Lee, K. A.; Liu, X.; Nautsch, A.; Patino, J.; Sahidullah, M.; Todisco, M.; Wang, X.; Yamagishi, J., ASVspoof 2021 Challenge - Logical Access Database, Zenodo, 2021.
  22. Delgado, H.; Evans, N.; Kinnunen, T.; Lee, K. A.; Liu, X.; Nautsch, A.; Patino, J.; Sahidullah, M.; Todisco, M.; Wang, X.; Yamagishi, J., ASVspoof 2021 Challenge - Physical Access Database, Zenodo, 2021.
  23. Delgado, H.; Evans, N.; Kinnunen, T.; Lee, K. A.; Liu, X.; Nautsch, A.; Patino, J.; Md Sahidullah; Todisco, M.; Wang, X.; Yamagishi, J., ASVspoof 2021 Challenge - Speech Deepfake Database, 2021.
  24. He, H.; Garcia, E. A., "Learning from Imbalanced Data," IEEE Transactions on Knowledge and Data Engineering, pp. 1263-1284, 2009.
  25. K. Ito and L. Johnson, "The LJ Speech Dataset," 2017. [Online]. Available: https://keithito.com/LJSpeech -Dataset/.
  26. Frank, J.; Schönherr, L., "WaveFake: A data set to facilitate audio DeepFake detection," Zenodo, 8 26, 2021. doi: 10.5281/zenodo.5642694.