Home > Published Issues > 2026 > Volume 17, No. 10, 2026 >
JAIT 2026 Vol.17(10): 1867-1876
doi: 10.12720/jait.17.10.1867-1876

Automatic Recognition of Speech Sounds Using Convolutional Neural Networks

Yasir Hashim 1,*, Mahmood Yashar 2, Hussein M. Ali 3, Abdulbari Yousif 2,
Mohd Abdul Rahim Khan 1, and Tarik A. Rashid 3
1. Department of Electronical Engineering and Computer Science, A’Sharqiyah University, Ibra, Oman
2. Department of Computer Engineering, Tishk International University, Erbil, Iraq
3. Computer Science and Engineering Department, University of Kurdistan Hewlêr, Erbil, Iraq
Email: yasir.naif@asu.edu.om (Y.H.); mahmood.yashar@tiu.edu.iq (M.Y.);
hussein.mohammedali@ukh.edu.krd (H.M.A.); abdulbaricmpe@gmail.com (A.Y.);
mohd.khan@asu.edu.om (M.A.R.K.); tarik.ahmed@ukh.edu.krd (T.A.R.)
*Corresponding author

Manuscript received April 7, 2026; revised May 18, 2026; accepted July 14, 2026; published October 9, 2026.

Abstract—Kurdish language has gained significant importance across various domains, yet it remains underrepresented in computational linguistics due to the scarcity of annotated datasets. This study develops and validates a robust Convolutional Neural Network (CNN)-based system for automatic recognition of Kurdish speech commands. A dedicated data-collection website was designed and deployed to gather over 4000 voice samples spanning 10 Kurdish commands, recorded by participants aged 7–64 across diverse environments, genders, and dialect backgrounds. The collected data (totaling 4527 samples) was partitioned into training (70%), validation (15%), and testing (15%) sets and preprocessed through low-pass filtering, amplitude normalization, and auditory spectrogram computation in MATLAB R2023b using the Audio Toolbox. The CNN architecture comprised three convolutional layers (3×3 filters, Rectified Linear Unit (ReLU) activation, 2×2 max-pooling with 0.25 dropout), followed by fully connected layers of 128 and 64 neurons, and a 10-class SoftMax output. The Adam optimizer (learning rate 0.001, batch size 32, 25 epochs) was used for training. The model achieved a validation accuracy of 98.17% and an independent test-set accuracy of 97.4%, with per-command recognition rates ranging from 91.5% to 100%. These results demonstrate the effectiveness of the proposed approach for Kurdish speech recognition and lay a foundation for future advancements in low-resource language processing.
 
Keywords—Convolutional Neural Network (CNN), Kurdish speech recognition, spectrogram, audio classification, low-resource language, deep learning
 
Cite: Yasir Hashim, Mahmood Yashar, Hussein M. Ali, Abdulbari Yousif, Mohd Abdul Rahim Khan, and Tarik A. Rashid, "Automatic Recognition of Speech Sounds Using Convolutional Neural Networks," Journal of Advances in Information Technology, Vol. 17, No. 10, pp. 1867-1876, 2026. doi: 10.12720/jait.17.10.1867-1876

Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions