Production-Oriented Speech AI Platform for Real-Time Multilingual Communication, Zero-Shot Voice Cloning, Live Interpretation, Synchronized Replay, Automated Quality Evaluation, and Reproducible Benchmarking.
LiveDub Enterprise is a production-oriented Speech AI platform for real-time multilingual communication. It combines LiveKit-based voice orchestration, streaming multilingual speech translation, zero-shot voice cloning, intelligent 9-signal provider routing, 1:1 synchronized session replay, automated speech quality evaluation (WER / CER / Speaker Similarity), and reproducible 2-layer verification pipelines into a unified platform architecture.
Designed for live interpretation, dubbing, accessibility, and high-stakes conference workflows across 12 languages.
- 🎙️ Real-Time Multilingual Translation: Low-latency 200ms chunked WebSockets pipeline across 12 languages (Hindi, English, Tamil, Telugu, Marathi, Kannada, Spanish, French, German, Japanese, Chinese, Arabic).
- 🎭 Zero-Shot Voice Cloning: Instant timbre transfer preserving speaker identity across target languages using Meta SeamlessM4T & CosyVoice2.
- ⚡ LiveKit Agent Orchestration: Stateful session management with asynchronous event-driven task distribution (
TranslationTask,VoiceCloneTask,EvaluationTask). - 🔀 Intelligent Provider Router: 9-signal score-based router evaluating provider health, streaming capabilities, voice clone support, latency, GPU acceleration, and queue depth.
- 📼 1:1 Synchronized Session Replay: Frame-by-frame session scrubber synchronizing audio playback, transcript advancement, pipeline glow, and telemetry metrics.
- 📊 Automated Speech Quality Evaluation: Real-time WER (Word Error Rate), CER (Character Error Rate), and PyTorch speaker similarity scoring.
- 🛡️ 2-Layer Reproducible Benchmarking: Separation of raw measurement data (
verification/raw_results.json) from dynamic markdown reporting with 95% confidence intervals and SHA256 integrity checksums.
LiveDub Enterprise Platform Architecture
┌────────────────────────────────────────────────────────────────────────┐
│ Next.js 14 Live Studio & Command Dashboard │
│ (Language Controls • Multi-Target Fan-out • Replay • Telemetry) │
└───────────────────────────────────┬────────────────────────────────────┘
│ REST / WebSockets / Control Frames
▼
┌────────────────────────────────────────────────────────────────────────┐
│ FastAPI Gateway & LiveKit Agent Supervisor │
│ (Session Context Engine • LanguageManager Single Source of Truth) │
└───────┬───────────────────────────┬───────────────────────────┬────────┘
│ │ │
▼ ▼ ▼
┌───────────────────┐ ┌───────────────────┐ ┌───────────────────┐
│ TranslationTask │ │ VoiceCloneTask │ │ EvaluationTask │
└─────────┬─────────┘ └─────────┬─────────┘ └─────────┬─────────┘
│ │ │
└───────────────────────────┼───────────────────────────┘
│
▼
┌───────────────────────────┐
│ Intelligent Score Router │
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ Speech AI Model Pool │
│ (SeamlessM4T • CosyVoice) │
└───────────────────────────┘
All percentiles computed dynamically from 152 live measured WebSockets audio iterations using numpy and statistics.quantiles (Deterministic Random Seed: 42).
| Sample Count | Min (ms) | Mean (ms) | P50 (Median) | P90 (ms) | P95 (ms) | P99 (ms) | Max (ms) | StdDev | 95% CI | E2E WS RTF | Model RTF |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 152 Chunks | 211.26 |
217.49 |
217.29 |
219.28 |
220.12 |
232.13 |
233.22 |
2.78 |
±0.44 |
1.087 |
0.239 |
RTF Methodology Notice:
- End-to-End WebSocket RTF (
$\mathbf{1.087}$ ): Total pipeline roundtrip over WebSocket transport ($217.49\text{ms} / 200\text{ms}$ chunk).- Pure Model Inference RTF (
$\mathbf{0.239}$ ): Pure GPU model execution on NVIDIA RTX 4090 ($47.8\text{ms} / 200\text{ms}$ chunk).
- Python: 3.10+
- Node.js: v18+ / v22+
- FFmpeg: Installed & on system PATH
git clone https://github.com/rajveer100704/LiveDub-Enterprise.git
cd LiveDub-Enterprise
# Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: .\venv\Scripts\activate
# Install dependencies
pip install -r requirements.txtpython -m uvicorn gateway.main:app --host 0.0.0.0 --port 8000cd frontend
npm install
npm run devOpen http://localhost:3000 in your browser to launch the Studio interface.
python scripts/verify_all.pyLiveDub-Enterprise/
├── config/ # Single source of truth (supported_languages.json)
├── gateway/ # FastAPI REST & WebSocket streaming server
├── agents/ # LiveKit Agent Supervisor & task runners
├── language/ # LanguageManager & LanguagePair dataclasses
├── session/ # SessionContext, SessionManager & SessionReplay
├── events/ # Decoupled EventBus pub-sub engine
├── providers/ # Model registry & provider integrations
├── registry/ # 9-signal Intelligent Provider Router
├── tasks/ # Modular workflow tasks (Translation, VoiceClone, Evaluation)
├── faults/ # Resilience & fault injection scenarios
├── verification/ # Reproducible verification outputs, SHA256, metadata & reports
├── frontend/ # Next.js 14 flagship Live Studio application
└── scripts/ # Automated verification & benchmark scripts
This project is licensed under the MIT License.