IPEM Group of Institutions
IPEM Journals

IPEM Journals/Computer Applications/Vol. 7

Audio Classification using Machine Learning

Shivanshi · Saakshi Joshi · Richa Singh

Vol. 7, Issue 1 (2022)Pages 10-16Published 2022-11-30

Abstract

The present work has been carried out on audio classification for outdoor events by making use of artificial neural network (ANN) The main aim of using ANN is to figure out a way of mapping the input with the corresponding output class present in the dataset. It is quite interesting to explore about how much machines are capable to recognize the audio and to determine what type of sound it is. In this research authors use audio clips taken from UrbanSound8K dataset are processedand converted to spectrograms that are further fed to the ANN classifier which identifies the type of sound based on thefrequency data captured in the spectrogram.

Keywords

  • Exploratory Data Analysis
  • Mel Frequency Cepstral Coefficient
  • Artificial Neural Network

Author affiliations

  • Shivanshi, School of Computing, DIT University, Dehradun, Uttarakhand, 248009
  • Saakshi Joshi, School of Computing, DIT University, Dehradun, Uttarakhand, 248009
  • Richa Singh, School of Computing, DIT University, Dehradun, Uttarakhand, 248009

How to cite

Keywords

References (25)

  1. 1.SoontornOraintara, Ying-Jui Chen Et.al. IEEE transactions on signal processing, IFFT, Vol. 50, No. 3, March 2002.
  2. 2.Kelly Wong, Journal of Undergraduate Research, Florida, Vol 2, Issue 11 - August 2011.
  3. 3.M.A. Anusuya and S.K. Katti. Speech recognition by machine: A review: International Journal of Computer Science and Information Security, Vol. 6, No. 3, 2009.
  4. 4.Sadaoki Furui, 50 years of progress in speech and Speaker Recognition Research, ECTI transactions on Computer and Information Technology, Vol.1. No.2 November 2005. .
  5. 5.D.R. Reddy, an approach to computer speech recognition by direct analysis of the Speech Wave, Tech. Report No.C549, Computer ScienceDept., Stanford Univ., September 1966. .
  6. 6.Simon Kinga and Joe Frankel, recognition, speech production knowledge in automatic speech recognition, Journal of Acoustic Society of America, 2006. .
  7. 7.Nathaniel Morgan and Herve Bourlard. Continuous speech recognition using multilayer perceptrons with Hidden Markov models. In IEEE International Conference on acoustics, speech and signal processing, 1990. .
  8. 8.Abdel-Rahman Mohamed, George E. Dahl, and Geoffrey E. Hinton. Deep belief networks for phone recognition. In neural information processing systems: Workshop on deep learning for speech recognition and related applications, 2009. .
  9. 9.Jen-Tzung Chien, linear regression base bayesia predictive classification for speech recognition, IEEE transactions on speech and audio processing, Vol. 11 No. 1, January 2003. .
  10. 10.Mohamed Afify and Olivier Siohan, sequential estimation with optimal forgetting for robust speech recognition, IEEE transactions on speech and audio processing, Vol. 12, No. 1, January 2004. .
  11. 11.Shinji Watanabe, variational bayesian estimation and clustering for speech recognition, IEEE transactions on speech and audio processing, Vol. 12, No. 4, July 2004. .
  12. 12.Mohamed Afify, Feng Liu, Hui Jiang, A new verification-based fast-match for large vocabulary continuous speech recognition, IEEE transactions on speech and audio processing, Vol. 13, No. 4, July 2005. .
  13. 13.Giuseppe Riccardi, active learning: Theory and applications to automatic speech recognition, IEEE transactions on speech and audio processing, Vol. 13, No. 4, July 2005. .
  14. 14.Frank Wessel and Hermann Ney , unsupervised training of acoustic models for large vocabulary continuous speech recognition, IEEE transactions on speech and audio processing, Vol. 13, No. 1, January 2005 . .
  15. 15.Mathias De-Wachter et.al., Template based continuous speech recognition, IEEE transactions on Audio, speech and Language processing, Vol.15, No.4, May 2007. .
  16. 16.J.W.Forgie and C.D.Forgie, results obtained from a vowel recognition computer program, J.A.S.A., 31(11),pp.1480-1489.1959. .
  17. 17.V.M. Velichko and N.G. Zagoruyko, automatic recognition of 200 words, Int.J.Man-Machine Studies, 2:223, June 1970. .
  18. 18.Michael Unser, Thierry Blu, IEEE Transactions on Signal Processing, Wavelet Theory Demystified, Vol. 51, No. 2, Feb’13.
  19. 19.Soundararajan et.al., Transforming Binary uncertainties for Robust Speech Recognition, IEEE Transactions on Audio, Speech and Language processing, Vol.15,No.6, July 2007. .
  20. 20.Ram Singh, proceedings of the NCC, spectral subtraction speech enhancement with RASTA filtering IIT-B 2012. .
  21. 21.Nitin Sawhney, situational awareness from environmental sounds, SIG, MIT Media Lab, June 13, 2013. .
  22. 22.Chadha, N., Gangwar, R.C., Bedi, R. 2015. current challenges and applications of speech recognition process using natural language processing: A survey, International Journal of Computer Applications (0975- 8887), 131(11), 28-31. .
  23. 23.Petkar, H. 2016. A review of challenges in “Automatic Speech Recognition”, International Journal of Computer Applications (0975-8887), 151(3), 23-29. .
  24. 24.Ramírez, J., Górriz, J.M., Segura, J.C. 2007. Voice activity detection, fundamentals and speech recognition systems robustness, University of Granada, Spain. .
  25. 25.Hirsch, H.G., Pearce, D. 2000. The aurora experimental framework for the performance evaluation of speech recognition system under noisy conditions, ASR-2000 Automatic speech recognition: Challenges for the new millennium Paris, France, 181-188.
Audio Classification using Machine Learning | IPEM Journals