Skip to content

About

- Speech-to-text processing with open-source models - Language translation using LLM-powered tools - Speech synthesis for real-time responses - Optimizing latency for seamless interaction - Deploying with open-source frameworks and APIs

Resources

Stars

1 star

Watchers

1 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

AI-Powered-Voice-to-Voice-Translator

  • Speech-to-text processing with open-source models - Language translation using LLM-powered tools - Speech synthesis for real-time responses - Optimizing latency for seamless interaction - Deploying with open-source frameworks and APIs

"# Call-Translation-ai"

Speech-to-Text, Translation & Text-to-Speech System

This Notebook transcribes speech using OpenAI Whisper, translates it from 60+ Languages to Sinhala using Google Translate, and converts the translated text into speech using gTTS (Google Text-to-Speech). The interface is built with Gradio for easy interaction.

Installation

Before running the Notebook file (.ipynb), you need to install the required dependencies.

Install Required Packages

Run the following command to install all necessary dependencies, pip install openai-whisper gradio scipy googletrans==4.0.0-rc1 gtts

Usage

Run the Jupyter Notebook.

Speak into the microphone.

The system will,

  • Convert speech to text (60+ Lanague ASR)
  • Translate the text to Sinhala
  • Convert the translated Sinhala text to speech (TTS)

The translated text will be displayed, and you can listen to the generated Sinhala speech.

Features

  • Speech-to-Text (STT) - Using OpenAI Whisper to transcribe speech(ASR,STT).
  • Translation - Using Google Translate API to translate recogised Lanague to Sinhala.
  • Text-to-Speech (TTS) - Converting Sinhala text into speech using gTTS.
  • User-Friendly Interface - Built using Gradio for easy interaction.

Notice

Whisper Model - This Notebook uses the small model. You can change it to base, medium, or large for different accuracy and performance levels.

Internet Requirement - Google Translate and gTTS require an active internet connection.

Audio Format - The system processes .wav files for transcription.

About

- Speech-to-text processing with open-source models - Language translation using LLM-powered tools - Speech synthesis for real-time responses - Optimizing latency for seamless interaction - Deploying with open-source frameworks and APIs

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages