- Speech-to-text processing with open-source models - Language translation using LLM-powered tools - Speech synthesis for real-time responses - Optimizing latency for seamless interaction - Deploying with open-source frameworks and APIs
"# Call-Translation-ai"
This Notebook transcribes speech using OpenAI Whisper, translates it from 60+ Languages to Sinhala using Google Translate, and converts the translated text into speech using gTTS (Google Text-to-Speech). The interface is built with Gradio for easy interaction.
Before running the Notebook file (.ipynb), you need to install the required dependencies.
Run the following command to install all necessary dependencies,
pip install openai-whisper gradio scipy googletrans==4.0.0-rc1 gtts
Run the Jupyter Notebook.
Speak into the microphone.
The system will,
- Convert speech to text (60+ Lanague ASR)
- Translate the text to Sinhala
- Convert the translated Sinhala text to speech (TTS)
The translated text will be displayed, and you can listen to the generated Sinhala speech.
- Speech-to-Text (STT) - Using OpenAI Whisper to transcribe speech(ASR,STT).
- Translation - Using Google Translate API to translate recogised Lanague to Sinhala.
- Text-to-Speech (TTS) - Converting Sinhala text into speech using gTTS.
- User-Friendly Interface - Built using Gradio for easy interaction.
Whisper Model - This Notebook uses the small model. You can change it to base, medium, or large for different accuracy and performance levels.
Internet Requirement - Google Translate and gTTS require an active internet connection.
Audio Format - The system processes .wav files for transcription.