Comparative Evaluation of Random Forest and 1D CNN Using MFCC Features for Embedded Voice-Command Recognition

Authors

  • Helmi Aprianto Undergraduate Program in Applied Electronics Engineering, Universitas Negeri Yogyakarta, Indonesia
  • Dessy Irmawati Undergraduate Program in Applied Electronics Engineering, Universitas Negeri Yogyakarta, Indonesia

DOI:

https://doi.org/10.21831/jeatech.v7i01.101430

Abstract

Comparative studies of Random Forest (RF) and Convolutional Neural Network (CNN) classifiers for voice-command recognition typically report only which model is more accurate, without explaining why the weaker model underperforms or assessing whether either model can be deployed on a resource-constrained microcontroller. This study addresses that gap by comparing RF and a one-dimensional CNN for Mel-Frequency Cepstral Coefficient (MFCC)-based voice classification, diagnosing the specific cause of RF's underperformance, and evaluating estimated deployment feasibility on an ESP32-S3 microcontroller. Using 15,400 augmented recordings from the Google Speech Commands dataset across eleven command classes, both models were evaluated on a dedicated test set. An RF classifier evaluated on time-averaged MFCC features reached only 49% test accuracy; the confusion matrix, feature-importance, principal component analysis, and centroid-distance analyses converged on one explanation: averaging across time erases the temporal structure of speech, collapsing acoustically similar classes into overlapping feature-space regions. Retaining the complete, un-averaged MFCC sequence raised the RF test accuracy to 73%, but the resulting model far exceeded a typical microcontroller's flash budget. Conversely, a CNN trained on the identical full-sequence representation reached 93% test accuracy. Based on hardware profiling estimates, this CNN requires only a kilobyte-scale flash memory footprint, replacing the massive multi-megabyte arrays required by RF. These results demonstrate that preserving temporal structure is critical for both classification accuracy and compact model size in embedded voice-command recognition. Future work will address current study limitations by incorporating silence and unknown background classes, and by ensuring strictly speaker-independent evaluation protocols for robust real-world deployment.

Author Biographies

Helmi Aprianto, Undergraduate Program in Applied Electronics Engineering, Universitas Negeri Yogyakarta, Indonesia

Undergraduate Program in Applied Electronics Engineering, Universitas Negeri Yogyakarta, Indonesia

Dessy Irmawati, Undergraduate Program in Applied Electronics Engineering, Universitas Negeri Yogyakarta, Indonesia

Undergraduate Program in Applied Electronics Engineering, Universitas Negeri Yogyakarta, Indonesia

Downloads

Published

2026-09-22

How to Cite

Aprianto, H., & Irmawati, D. (2026). Comparative Evaluation of Random Forest and 1D CNN Using MFCC Features for Embedded Voice-Command Recognition. Journal of Engineering and Applied Technology, 7(01), 44–59. https://doi.org/10.21831/jeatech.v7i01.101430

Issue

Section

Articles