Home
World Journal of Advanced Engineering Technology and Sciences
International, Peer reviewed, Referred, Open access | ISSN Approved Journal

Main navigation

  • Home
    • Journal Information
    • Abstracting and Indexing
    • Editorial Board Members
    • Reviewer Panel
    • Journal Policies
    • WJAETS CrossMark Policy
    • Publication Ethics
    • Instructions for Authors
    • Article processing fee
    • Track Manuscript Status
    • Get Publication Certificate
    • Issue in Progress
    • Current Issue
    • Past Issues
    • Become a Reviewer panel member
    • Join as Editorial Board Member
  • Contact us
  • Downloads

ISSN: 2582-8266 (Online)  || UGC Compliant Journal || Google Indexed || Impact Factor: 9.48 || Crossref DOI

Fast Publication within 2 days || Low Article Processing charges || Peer reviewed and Referred Journal

Research and review articles are invited for publication in Volume 20, Issue 3 (September 2026).... Submit articles

A Multimodal Deep Learning Model for Emotion Recognition from Video, Audio and Text Data

Breadcrumb

  • Home
  • A Multimodal Deep Learning Model for Emotion Recognition from Video, Audio and Text Data

Ogochukwu Patience Okechukwu 1, * and Godson Nnaeto Okechukwu 2

1 Department of Computer Science, Faculty of Physical Science, Nnamdi Azikiwe University, Awka, Nigeria.
2 Department of Electronic/Computer Engineering, Faculty of Engineering, Nnamdi Azikiwe University, Awka, Nigeria.

Research Article

 

World Journal of Advanced Engineering Technology and Sciences, 2026, 19(02), 121-135

Article DOI: 10.30574/wjaets.2026.19.2.0257

DOI url:https://doi.org/10.30574/wjaets.2026.19.2.0257

Received on 02 April 2026; revised on 11 May 2026; accepted on 13 May 2026

This study presents the development of a multimodal deep learning model for emotion recognition using video, audio, and text data extracted from a video data. The proposed system employs a structured pipeline that begins with video input, from which audio is extracted and transcribed into text using the Whisper model. Emotion recognition is performed across three modalities: facial expressions analysed using DeepFace, vocal features classified using a Long Short-Term Memory (LSTM) network, and textual features processed using a Feedforward Neural Network (FNN). Results are visualized in bar chart plotted using Seaborn and integrated into an interactive user interface built with Flet. This multimodal approach enhances accuracy and robustness over unimodal systems by leveraging complementary information from each data stream. The model has broad applications in affective computing, human-computer interaction, surveillance, and sentiment-aware technologies.

Multimodal; DeepFace; LSTM; FNN; Deep Learning; Emotion

https://wjaets.com/sites/default/files/fulltext_pdf/WJAETS-2026-0257.pdf

Get Your e Certificate of Publication using below link

Download Certificate

Preview Article PDF

Ogochukwu Patience Okechukwu and Godson Nnaeto Okechukwu. A Multimodal Deep Learning Model for Emotion Recognition from Video, Audio and Text Data. World Journal of Advanced Engineering Technology and Sciences, 2026, 19(02), 121-135. Article DOI: https://doi.org/10.30574/wjaets.2026.19.2.0257

Get Certificates

Get Publication Certificate

Download LoA

Check Corssref DOI details

Issue details

Issue Cover Page

Editorial Board

Table of content


Copyright © Author(s). All rights reserved. This article is published under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits use, sharing, adaptation, distribution, and reproduction in any medium or format, as long as appropriate credit is given to the original author(s) and source, a link to the license is provided, and any changes made are indicated.


Copyright © 2026 World Journal of Advanced Engineering Technology and Sciences

Developed & Designed by VS Infosolution