🎭 Moody.AI - Advanced Multimodal Sentiment Analysis
A production-ready AI application that performs real-time emotion analysis from video content using state-of-the-art deep learning models across three modalities: computer vision, audio processing, and natural language processing.
🧠 AI Architecture & Models
Computer Vision Pipeline
Face Detection : CVLib + OpenCV for robust facial region extraction
Vision Transformer : DINOv2 ViT-B/14 (518×518 input) for visual feature extraction
Preprocessing : ImageNet normalization, temporal sampling at ~1 FPS i.e 1 image per second of the Video Uploaded
Audio Processing Pipeline
Speech-to-Text : OpenAI Whisper (base/small/tiny models) with multilingual support
Audio Features : Finetuned-audeering/wav2vec2-large-robust-12-ft-emotion-msp-dim (Freeze = 8 , Unfreeze = 4) for emotional audio representations with the help of RAVDESS and MELD Audio Dataset.
Audio Format : 16kHz WAV extraction via MoviePy
Natural Language Processing
Text Embeddings : Fine-Tuned DistilBERT-base-uncased (10:2) with hidden state averaging
Sentiment Analysis : Transformer-based sequence classification
Language Support : Multilingual with optional language hints
Multimodal Fusion
Architecture : Custom TrimodalFusionModel with cross-attention mechanisms
Integration : Attention-based fusion of vision, audio, and text features
Output : 5-class emotion classification (anger, joy, melancholy, neutral, surprise)
📊 Model Performance
Vision Accuracy : Optimized for facial emotion recognition. For Vison part of MELD ~ 23.5%
Audio Accuracy : Leverages Wav2Vec2's pre-trained emotional representations . For Audio MELD ~ 38.18%
Text Accuracy : DistilBERT fine-tuned for sentiment classification. For TEXT ~ 55.5%
Fusion Model : Trained on multimodal emotion datasets with cross-attention. After Tri-Modal Fusion and Additional Layers ~ 61.05% Acc and 62% Recall
🎨 User Interface
Framework : Streamlit with custom CSS and animations
Design : Glass morphism UI with animated gradient backgrounds
Features : Real-time processing indicators, confidence scores, probability distributions
Responsive : Works on desktop and mobile browsers
🚀 Deployment & MLOps
Containerization : Docker with optimized dependency management
Base Image : Python 3.11 with CUDA support
Dependencies : Resolved complex ML library conflicts (TensorFlow, PyTorch, OpenCV)
Memory Management : CUDA optimization and automatic model cleanup
Build Size : ~4.1 GB with all ML dependencies
⚡ Quick Start
Run the application
docker run -p 8501:8501 rishab27279/moody-ai
Access at http://localhost:8501
🔧 Technical Specifications
Processing Modes : Lightweight, Balanced, High Fidelity
Input Formats : MP4, MOV, MKV (up to 200MB)
Output : Emotion classification with confidence scores
Hardware : CPU/GPU compatible with automatic device detection
🌟 Features
✅ Real-time multimodal emotion analysis
✅ Beautiful animated user interface
✅ Multilingual speech recognition
✅ Production-ready Docker deployment
✅ Memory-efficient processing
✅ Cross-platform compatibility
🏗️ Built With
Deep Learning : PyTorch, TensorFlow, Transformers, timm
Computer Vision : OpenCV, CVLib, DINOv2
Audio Processing : Librosa, Whisper, Wav2Vec2
Web Framework : Streamlit
Deployment : Docker, WSL2 compatible
📈 Use Cases
Academic research in multimodal AI
Emotion analysis for video content
Educational demonstrations of ML pipelines
Baseline for multimodal sentiment analysis projects
Built with ❤️ for the AI community. Contributions and feedback welcome!