Comparative Analysis of MFCC and MFE Feature Extraction Methods for CNN-Based Railway Sound Classification

Authors

  • Nadaval Hagga Politeknik Negeri Malang
  • Sapto Wibowo Politeknik Negeri Malang
  • Moechammad Sarosa Politeknik Negeri Malang

DOI:

https://doi.org/10.33795/jartel.v16i3.10281

Keywords:

Railway Sound Classification, MFCC, MFE, CNN, Raspberry Pi 5

Abstract

Environmental sound classification is one of the artificial intelligence fields widely applied in embedded system-based monitoring applications. This study compares two audio feature extraction methods, namely Mel-Frequency Cepstral Coefficients (MFCC) and Mel-Filterbank Energy (MFE), for railway sound classification using a Convolutional Neural Network (CNN) model. The dataset consisted of 153 audio samples collected directly from real-world environments, comprising 80 railway sound samples and 73 environmental noise samples. The dataset was divided into 122 training samples and 31 testing samples. Feature extraction and model training were performed using the Edge Impulse platform, and the trained models were subsequently implemented on a Raspberry Pi 5. The training results showed that MFCC achieved 96.7% accuracy, 97% precision, 97% recall, and 97% F1-score, while MFE achieved 98.4% accuracy, 98% precision, 98% recall, and 98% F1-score. During implementation on the Raspberry Pi 5, the MFCC-based model achieved 96.77% accuracy, while the MFE-based model achieved 100% accuracy. However, the MFCC-based model demonstrated lower computational latency, with an average inference time of 12.317 ms, compared to 23.44 ms for the MFE-based model. These results indicate that MFE provides higher classification performance, whereas MFCC provides better computational efficiency in terms of inference time, demonstrating a trade-off between classification performance and computational efficiency for railway sound classification on embedded platforms.

References

A. Guzhov, F. Raue, J. Hees, and A. Dengel, “ESResNet: Environmental Sound Classification Based on Visual Domain Models,” in Proc. Int. Conf. Pattern Recognition (ICPR), 2020, pp. 492–499.

J. Sharma, O. C. Granmo, and M. Goodwin, “Environment Sound Classification Using Multiple Feature Channels and Attention Based Deep Convolutional Neural Network,” arXiv preprint arXiv:1908.11219, 2019.

S. Hershey et al., “CNN Architectures for Large-Scale Audio Classification,” in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 131–135

D. Stowell and M. D. Plumbley, “Automatic Large-Scale Classification of Bird Sounds is Strongly Improved by Unsupervised Feature Learning,” PeerJ, vol. 2, p. e488, 2014.

E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “FSD50K: An Open Dataset of Human-Labeled Sound Events,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 30, pp. 829–852, 2022.

B. Zhu, K. Xu, D. Wang, L. Zhang, B. Li, and Y. Peng, “Environmental Sound Classification Based on Multi-Temporal Resolution Convolutional Neural Network Combining with Multi-Level Features,” IEEE Access, vol. 6, pp. 51412–51420, 2018

J. Salamon, C. Jacoby, and J. P. Bello, “Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification,” IEEE Signal Process. Lett., vol. 24, no. 3, pp. 279–283, 2017.

Y. Tokozume, Y. Ushiku, and T. Harada, “Learning from Between-Class Examples for Deep Sound Recognition,” in Proc. Int. Conf. Learning Representations (ICLR), 2018.

M. Cartwright, B. Cramer, J. Salamon, O. Nov, and J. P. Bello, “SONYC Urban Sound Tagging (SONYC-UST): A Multilabel Dataset from an Urban Acoustic Sensor Network,” in Proc. IEEE Workshop Applications Signal Processing to Audio and Acoustics (WASPAA), 2019, pp. 35–39.

Y. Su, K. Zhang, J. Wang, and K. Madani, “Environmental Sound Classification Using a Two-Stream CNN Based on Decision-Level Fusion,” Sensors, vol. 19, no. 7, p. 1733, 2019.

S. M. Abdoli, P. Cardinal, and A. L. Koerich, “End-to-End Environmental Sound Classification Using a 1D Convolutional Neural Network,” Expert Syst. Appl., vol. 136, pp. 252–263, 2019.

M. A. Ferrag, L. Maglaras, and A. Ahmim, “Audio Classification and Sound Event Detection: A Survey,” IEEE Access, vol. 11, pp. 13487–13525, 2023

N. Zeghidour, O. Teboul, F. de Chaumont Quitry, and M. Tagliasacchi, “LEAF: A Learnable Frontend for Audio Classification,” ICLR, 2021.

R. Ardiansyah, B. Bastiar, Adzikirani, D. Marya, and A. Novianti, “Optimasi Penerjemahan Bahasa Asing dengan Teknologi IoT pada Kelas Internasional Politeknik Negeri Malang,” Jurnal ELTEK, vol. 23, no. 1, pp. 1–8, Apr. 2025

D. Chicco and G. Jurman, “The Advantages of the Matthews Correlation Coefficient (MCC) over F1 Score and Accuracy in Binary Classification Evaluation,” BMC Genomics, vol. 21, no. 1, pp. 1–13, 2020.

S. Davis and P. Mermelstein, “Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences,” IEEE Trans. Acoust., Speech, Signal Process., vol. 28, no. 4, pp. 357–366, 1980

M. Ravanelli and Y. Bengio, “Speaker Recognition from Raw Waveform with SincNet,” in Proc. IEEE Spoken Language Technology Workshop (SLT), 2018, pp. 1021–1028.

K. J. Piczak, “Environmental Sound Classification with Convolutional Neural Networks,” in Proc. IEEE Int. Workshop Machine Learning for Signal Processing (MLSP), 2015, pp. 1–6.

R. A. Asmara, M. Ridwan, G. B. P., and A. N. Handayani, “Pengembangan Sistem Face Recognition Menggunakan Cloud Service, Raspberry Pi dan Convolutional Neural Network (CNN),” Jurnal ELTEK, pp. 95–102, 2022.

D. Aldiani, G. Dwilestari, H. Susana, R. Hamonangan, and D. Pratama, “Implementasi Algoritma CNN dalam Sistem Absensi Berbasis Pengenalan Wajah,” Jurnal ELTEK, pp. 197–202, 2024.

V. Tarigan, A. Yusupa, and R. Syahputra, “Optimasi Deteksi Penyakit Alzheimer dengan Convolutional Neural Network (CNN) dan Support Vector Machine (SVM) untuk Klasifikasi Tingkat Demensia,” Jurnal ELTEK, pp. 453–460, 2025.

Downloads

Published

30-09-2026

How to Cite

Hagga, N., Wibowo, S., & Sarosa, M. (2026). Comparative Analysis of MFCC and MFE Feature Extraction Methods for CNN-Based Railway Sound Classification. JURNAL JARTEL: Jurnal Jaringan Telekomunikasi, 16(3), 257–267. https://doi.org/10.33795/jartel.v16i3.10281