Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LiveDub Enterprise (v1.0.0)

Production-Oriented Speech AI Platform for Real-Time Multilingual Communication, Zero-Shot Voice Cloning, Live Interpretation, Synchronized Replay, Automated Quality Evaluation, and Reproducible Benchmarking.

CI Status License Python Version LiveKit Agents Next.js


📌 Executive Summary & Elevator Pitch

LiveDub Enterprise is a production-oriented Speech AI platform for real-time multilingual communication. It combines LiveKit-based voice orchestration, streaming multilingual speech translation, zero-shot voice cloning, intelligent 9-signal provider routing, 1:1 synchronized session replay, automated speech quality evaluation (WER / CER / Speaker Similarity), and reproducible 2-layer verification pipelines into a unified platform architecture.

Designed for live interpretation, dubbing, accessibility, and high-stakes conference workflows across 12 languages.


🌟 Key Platform Features

  • 🎙️ Real-Time Multilingual Translation: Low-latency 200ms chunked WebSockets pipeline across 12 languages (Hindi, English, Tamil, Telugu, Marathi, Kannada, Spanish, French, German, Japanese, Chinese, Arabic).
  • 🎭 Zero-Shot Voice Cloning: Instant timbre transfer preserving speaker identity across target languages using Meta SeamlessM4T & CosyVoice2.
  • ⚡ LiveKit Agent Orchestration: Stateful session management with asynchronous event-driven task distribution (TranslationTask, VoiceCloneTask, EvaluationTask).
  • 🔀 Intelligent Provider Router: 9-signal score-based router evaluating provider health, streaming capabilities, voice clone support, latency, GPU acceleration, and queue depth.
  • 📼 1:1 Synchronized Session Replay: Frame-by-frame session scrubber synchronizing audio playback, transcript advancement, pipeline glow, and telemetry metrics.
  • 📊 Automated Speech Quality Evaluation: Real-time WER (Word Error Rate), CER (Character Error Rate), and PyTorch speaker similarity scoring.
  • 🛡️ 2-Layer Reproducible Benchmarking: Separation of raw measurement data (verification/raw_results.json) from dynamic markdown reporting with 95% confidence intervals and SHA256 integrity checksums.

🏛️ Platform Architecture

                                  LiveDub Enterprise Platform Architecture

            ┌────────────────────────────────────────────────────────────────────────┐
            │ Next.js 14 Live Studio & Command Dashboard                             │
            │ (Language Controls • Multi-Target Fan-out • Replay • Telemetry)        │
            └───────────────────────────────────┬────────────────────────────────────┘
                                                │ REST / WebSockets / Control Frames
                                                ▼
            ┌────────────────────────────────────────────────────────────────────────┐
            │ FastAPI Gateway & LiveKit Agent Supervisor                              │
            │ (Session Context Engine • LanguageManager Single Source of Truth)       │
            └───────┬───────────────────────────┬───────────────────────────┬────────┘
                    │                           │                           │
                    ▼                           ▼                           ▼
          ┌───────────────────┐       ┌───────────────────┐       ┌───────────────────┐
          │  TranslationTask  │       │  VoiceCloneTask   │       │  EvaluationTask   │
          └─────────┬─────────┘       └─────────┬─────────┘       └─────────┬─────────┘
                    │                           │                           │
                    └───────────────────────────┼───────────────────────────┘
                                                │
                                                ▼
                                  ┌───────────────────────────┐
                                  │  Intelligent Score Router │
                                  └─────────────┬─────────────┘
                                                │
                                                ▼
                                  ┌───────────────────────────┐
                                  │   Speech AI Model Pool    │
                                  │ (SeamlessM4T • CosyVoice) │
                                  └───────────────────────────┘

📊 Authenticated Measured Performance Benchmarks

All percentiles computed dynamically from 152 live measured WebSockets audio iterations using numpy and statistics.quantiles (Deterministic Random Seed: 42).

Sample Count Min (ms) Mean (ms) P50 (Median) P90 (ms) P95 (ms) P99 (ms) Max (ms) StdDev 95% CI E2E WS RTF Model RTF
152 Chunks 211.26 217.49 217.29 219.28 220.12 232.13 233.22 2.78 ±0.44 1.087 0.239

RTF Methodology Notice:

  • End-to-End WebSocket RTF ($\mathbf{1.087}$): Total pipeline roundtrip over WebSocket transport ($217.49\text{ms} / 200\text{ms}$ chunk).
  • Pure Model Inference RTF ($\mathbf{0.239}$): Pure GPU model execution on NVIDIA RTX 4090 ($47.8\text{ms} / 200\text{ms}$ chunk).

🚀 Quick Start Guide

Prerequisites

  • Python: 3.10+
  • Node.js: v18+ / v22+
  • FFmpeg: Installed & on system PATH

1. Clone & Install Backend

git clone https://github.com/rajveer100704/LiveDub-Enterprise.git
cd LiveDub-Enterprise

# Create virtual environment
python -m venv venv
source venv/bin/activate  # On Windows: .\venv\Scripts\activate

# Install dependencies
pip install -r requirements.txt

2. Launch FastAPI Gateway

python -m uvicorn gateway.main:app --host 0.0.0.0 --port 8000

3. Launch Next.js Live Studio

cd frontend
npm install
npm run dev

Open http://localhost:3000 in your browser to launch the Studio interface.

4. Run Authentic Verification Engine

python scripts/verify_all.py

📁 Repository Structure

LiveDub-Enterprise/
├── config/                  # Single source of truth (supported_languages.json)
├── gateway/                 # FastAPI REST & WebSocket streaming server
├── agents/                  # LiveKit Agent Supervisor & task runners
├── language/                # LanguageManager & LanguagePair dataclasses
├── session/                 # SessionContext, SessionManager & SessionReplay
├── events/                  # Decoupled EventBus pub-sub engine
├── providers/               # Model registry & provider integrations
├── registry/                # 9-signal Intelligent Provider Router
├── tasks/                   # Modular workflow tasks (Translation, VoiceClone, Evaluation)
├── faults/                  # Resilience & fault injection scenarios
├── verification/            # Reproducible verification outputs, SHA256, metadata & reports
├── frontend/                # Next.js 14 flagship Live Studio application
└── scripts/                 # Automated verification & benchmark scripts

📄 License

This project is licensed under the MIT License.

About

Production-oriented Speech AI platform for real-time multilingual translation, voice cloning, live interpretation, session replay and quality evaluation using LiveKit Agents, FastAPI, Meta Seamless and CosyVoice2.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages