1 Department of Computer Science, Faculty of Physical Science, Nnamdi Azikiwe University, Awka, Nigeria.
2 Department of Electronic/Computer Engineering, Faculty of Engineering, Nnamdi Azikiwe University, Awka, Nigeria.
World Journal of Advanced Engineering Technology and Sciences, 2026, 19(02), 121-135
Article DOI: 10.30574/wjaets.2026.19.2.0257
Received on 02 April 2026; revised on 11 May 2026; accepted on 13 May 2026
This study presents the development of a multimodal deep learning model for emotion recognition using video, audio, and text data extracted from a video data. The proposed system employs a structured pipeline that begins with video input, from which audio is extracted and transcribed into text using the Whisper model. Emotion recognition is performed across three modalities: facial expressions analysed using DeepFace, vocal features classified using a Long Short-Term Memory (LSTM) network, and textual features processed using a Feedforward Neural Network (FNN). Results are visualized in bar chart plotted using Seaborn and integrated into an interactive user interface built with Flet. This multimodal approach enhances accuracy and robustness over unimodal systems by leveraging complementary information from each data stream. The model has broad applications in affective computing, human-computer interaction, surveillance, and sentiment-aware technologies.
Multimodal; DeepFace; LSTM; FNN; Deep Learning; Emotion
Get Your e Certificate of Publication using below link
Preview Article PDF
Ogochukwu Patience Okechukwu and Godson Nnaeto Okechukwu. A Multimodal Deep Learning Model for Emotion Recognition from Video, Audio and Text Data. World Journal of Advanced Engineering Technology and Sciences, 2026, 19(02), 121-135. Article DOI: https://doi.org/10.30574/wjaets.2026.19.2.0257