You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Audio tagging is the process of inferring descriptive labels from audio clips (Multi label classification task). This repository contains exploratory code/scripts for audio preprocessing and model fitting for the task of audio tagging and its applications.
An intelligent speech recognition system that combines OpenAI's Whisper for accurate transcription with dual emotion detection models. Analyzes both audio characteristics (tone, pitch, intensity) and textual content to provide comprehensive emotional context alongside transcriptions.
Interactive AI platform for public speaking and impromptu speech practice with multi-metric audio evaluation, real-time feedback, and personalized AI coaching.
AI system that analyzes urban soundscapes to optimize city planning, reduce noise pollution, and enhance acoustic environments using audio deep learning.
This Python script is for a voice interface chatbot named Jervis. It uses OpenAI's GPT-3.5-turbo-instruct model to respond to user input. The chatbot responds by Elevenlabs Voices. Conversation are saved to MongoDB, and MP3 file local and can be emailed if needed.
This project automates audio processing by removing silence, transcribing speech to text, and storing the output in an SQLite database. It supports multiple audio formats and leverages Google Speech Recognition for high accuracy.
Conversation-Mixer-Tool: A Python utility to merge two audio files (caller and receiver) into a seamless, conversation-like output. Features include speech detection, bandpass filtering, noise reduction, and smooth audio transitions.
A deep learning speech emotion recognition system using CNN-LSTM, MFCC feature extraction, and audio augmentation to analyze emotional patterns and explore their potential for early depression screening.