AI-Powered Speech Analysis Tool for Speech Therapists
DOI:
https://doi.org/10.21467/proceedings.7.5.4Keywords:
Speech analysis, Speech disorders, AI healthcare toolAbstract
Speech disorders present a considerable challenge within the realms of clinical diagnosis and therapeutic intervention, thereby necessitating the development of practical tools for the analysis and monitoring of speech patterns. This mini project, entitled "AI-Powered Speech Analysis Tool for Speech Therapists," was developed to integrate advanced technology into the diagnostic process and provide an innovative solution for therapists. The system enabled therapists to record patients' speech, analyze essential acoustic features such as pitch, speech rate, volume, and articulation patterns, and visualize the results through intuitive graphical representations, including spectrograms and pitch contours. It further facilitated comparative analyses against normative speech patterns, identified anomalies, and suggested potential speech disorders based on the derived analyses. Developed using the Python programming language and leveraging libraries such as librosa for audio processing, matplotlib for data visualization, and SQLite for data storage, the tool ensured broad accessibility and expandability. Diagnostic suggestions, derived from either rule-based systems or machine learning models, augmented its utility, while strict adherence to data privacy regulations safeguarded patient confidentiality. The application holds the potential to be extended into a web-based interface employing frameworks like Flask or Django, thereby enhancing its accessibility for therapists. This project demonstrates technical proficiency in speech processing and data visualization and contributes to the field of speech-language pathology by improving the accuracy and efficiency of speech disorder diagnosis. By facilitating the monitoring of patient progress and enhancing therapeutic outcomes, this tool presents opportunities for further advancements in AI-powered healthcare applications.
References
[1] P. F. Jr., "Emergence and prevalence of persistent and residual speech errors," Seminars in Speech and Language, vol. 36, no. 4, pp. 217-223, November 2015.
[2] D. A. M. e. al., "Classifying Speech Disorders Using Voice Signals and Machine Learning," in 2025 22nd International Learning and Technology Conference (L&T), 2025.
[3] B. B. Lubker, "Epidemiology: An essential science for speech-language pathology and audiology," J. Commun. Disord., vol. 30, no. 4, pp. 251-267, 1997.
[4] G. Ying, "Research on English Speech Data Analysis Algorithm Based on Deep Learning," in 2024 International Conference on Language Technology and Digital Humanities (LTDH), Bhubaneswar, 2024.
[5] A. R. K. e. al., "Voice Sentiment Analysis System Using Librosa," in 2024 4th International Conference on Ubiquitous Computing and Intelligent Information Systems (ICUIS), 2024.
[6] A. Ravindran, Django Design Patterns and Best Practices: Industry-standard web development techniques and solutions using Python, Packt Publishing Ltd., 2018.
[7] J. C. Y. H. a. I. C. J. Benesty, Noise Reduction in Speech Processing, vol. 2, Springer Science & Business Media, 2009.
[8] N. Dave, "Feature extraction methods LPC, PLP and MFCC in speech recognition," Int. J. Advance Res. Eng. Technol., vol. 1, no. 6, pp. 1-4, 2013.
[9] A. V. Oppenheim, "Speech spectrograms using the fast Fourier transform," IEEE Spectrum, vol. 7, no. 8, pp. 57-62, August 1970.
[10] A. d. Cheveigné and H. Kawahara, "YIN, a fundamental frequency estimator for speech and music," Journal of the Acoustical Society of America, vol. 111, no. 4, pp. 1917-1930, April 2002.
[11] G. E. A. P. A. Batista, R. C. Prati and M. C. Monard, "A study of the behavior of several methods for balancing machine learning training data," ACM SIGKDD, vol. 6, no. 1, pp. 20-29, June 2004.
[12] P. Gupta, "Dysarthria and Non-Dysarthria Speech Dataset," Kaggle, 2023.
[13] F. R. L. K. M. S. S. P. A. N. J. Y. P. M. v. L. D. J. O. J. R. Tobin, "The TORGO Database of Dysarthric Speech," University of Toronto, 2010.
[14] S. V. Thambi, K. T. Sreekumar, C. S. Kumar and P. R. Raj, "Random forest algorithm for improving the performance of speech/non-speech detection," in 2014 First International Conference on Computational Systems and Communications (ICCSC), 2014.
[15] K. Kimiafar, M. Sarbaz, D. Sobhani-Rad, A. S. Mousavi, M. R. Mehneh, F. D. Kemmak, S. F. M. Baigi and M. Esmaeili, "A comparative study of minimum data set of speech therapy: A systematic literature review," Frontiers in Health Informatics, vol. 12, 2023.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.