amanbasu / speech-emotion-recognition Goto Github PK

View Code? Open in Web Editor NEW

123.0 4.0 38.0 2.6 MB

Detecting emotions using MFCC features of human speech using Deep Learning

License: GNU General Public License v3.0

Jupyter Notebook 96.60% Python 3.40%

tensorflow deep-learning rnn mfcc speech-recognition emotion-recognition emotion

speech-emotion-recognition's Introduction

Recognising Human Emotions From Raw Audio

Collaborator: Aman Agarwal, Aditya Mishra

In this project we will use Mel frequency cepstral coefficients (MFCC) to train a recurrent neural network (LSTM) and classify human emotions into happy, sad, angry, frustrated, sad, neutral and fear categories.

The dataset used is The Interactive Emotional Dyadic Motion Capture (IEMOCAP) collected by University of Southern California

the link for the same can be found here

The dataset

The IEMOCAP database consists of 10 emotions. We selected the major 6 emotions viz. angry, neutral, frustrated, sad, excited and happy, in our training set. Features extracted from the raw audio of all sessions were saved along with their length and emotion. We used the first 20 mfcc coefficients as the feature vector, the process can be found in notebook

To convert data into a consistent shape we have applied Bucket Padding. The data is first sorted according to their sequence lengths and then divided into a specific number of buckets. The length of data thus divided is in close range of each other which eliminates extra padding. This method is used in Bucket Iterator which is used to get the batch if desired examples.

For selecting a batch, a bucket is chosen at random containing sorted data, out of that bucket contiguous examples equal to the batch size are chosen. The examples are padded to the shape of maximum sequence length and then shuffled. This gives the desired batch. the code for bucket iterator is taken from R2RT

Model

We used two layers of Bidirectional LSTM followed by attention in the last layer. The batch size was kept as 128 with the learning rate of 1e-4.

Results

The model was trained for 500 epochs and after which the curve almost reached a plateau. The model showed overfitting when the dropout was not used. We then applied a dropout of keep probability 0.8 between the last LSTM layer and the output layer.

Adding dropout reduced the overfitting of the model and increased its overall accuracy. The model showed an unweighted accuracy across six emotions of 45% with the validation accuracy of 42%.

Dropout of 0.2	No Dropout

Tensorflow model

Tensorflow implementation of the model has been added. The repository contains two files, speech_emotion_gpu to run the model on gpu and speech_emotion_gpu_multi which makes the file run parallelly on multiple gpus.

Input data for model can be downloaded from this link.

It consists of the following features: F0 (pitch), voice probability, zero-crossing rate, 12-dimensional Mel-frequency cepstral coefficients (MFCC) with log energy, and their first time derivatives. The features have been taken from this paper.

speech-emotion-recognition's People

Contributors

Stargazers

Watchers

speech-emotion-recognition's Issues

FUNCTION OF BUCKET ITERATOR

I want to know what is the main function of bucket iterator and what are parameters and return value.

misssing files Ses01F_script01_2_F008.wav and microsoft_32_features.pkl

Dear sir,

I appropriate your work, i'm student at GMR Institute of Technology,Rajam. As a part of the Project, i am working (Speech to emotion recognition using RNN). For reference i have seen your github. (https://github.com/amanbasu/speech-emotion-recognition.git).

But there are some files are missing in the repository.

Could you please share those data with me?

Ses01F_script01_2_F008.wav
microsoft_32_features.pkl

i couldn't understand the

"path_to_database/IEMOCAP_database/Session{}/wav//.wav".

because of this i couldn't able to load the data...Please help me in this regard

Please respond to the mail as early as possible

Thank you.

Where is the speech_emotion_data.pkl?

Hello Aman:

Thank you for enjoying your code.I am trying to follow your code.But I don't find the "speech_emotion_data.pkl" in the file of "speech_emotion_gpu.py".So please could you provide the speech_emotion_data.pkl ? Thanks very much.

Best wishes!

zero_crosses should be '/320' not '/0.02', shouldn't it?

In extract_features.py, the following code should be /320 not /0.02, shouldn't it?

zero_crosses = np.nonzero(np.diff(sig[start:end] > 0))[0].shape[0]/0.02 # zero crosses

↓ modify

zero_crosses = np.nonzero(np.diff(sig[start:end] > 0))[0].shape[0]/ 320 # zero crosses

※ Zero Crossing Rate : The rate of sign-changes of the signal during the duration of a particular frame. [1]

ref.[1]: https://github.com/tyiannak/pyAudioAnalysis/wiki/3.-Feature-Extraction

Link for speech_emotion_data.pkl

Hi,
Could you please provide me the link for speech_emotion_data.pkl file that you used in speech_emotion_gpu.py python file.
And also if possible could you upload the requirements.txt file to know the versions of the packages you used since I was facing so many errors with version changes.

Thanks in advance!

Thanks in advance, appreciate your help.