A New End-to-end Framework for Evaluating Voice Agents
-
Updated
Sep 29, 2026 - Python
A New End-to-end Framework for Evaluating Voice Agents
🇺🇦 Open Source Ukrainian Text-to-Speech datasets
A Docker-based OpenAI-compatible Text-to-Speech API server powered by Kyutai's TTS models with GPU acceleration support.
open-source speech AI platform for organizations that cannot send sensitive conversations to a third party
A unified benchmarking framework for evaluating Voice AI agents across conversational quality, audio realism, latency metrics, and safety guardrails with scalable multi-language stress testing.
MLX Porting Toolkit — an agent-guided, evidence-gated pipeline (scaffold → convert → parity → benchmark) plus a portable skill for porting PyTorch/Hugging Face models to Apple MLX.
An open, modular framework for building and evaluating low-resource speech-AI pipelines
Open source AI voice calling agent for Twilio phone calls, built with FastAPI, Google ADK, and Gemini Live API
中文 ASR 评测工具箱 · micro-CER 对比 FunASR/Whisper/llama.cpp · 一条命令出报告 · 自带迷你测试集 · Mandarin ASR benchmark toolkit
Production-oriented Speech AI platform for real-time multilingual translation, voice cloning, live interpretation, session replay and quality evaluation using LiveKit Agents, FastAPI, Meta Seamless and CosyVoice2.
Lucida Harness is a lightweight test harness and local Web UI for evaluating background removal models and alpha matting algorithms.
Multilingual AI speech studio for Text-to-Speech, Speech-to-Text, Voice Cloning, and Audio Enhancement in Uzbek, English, and Korean.
Ask questions about uploaded documents by voice, hear multilingual replies, and extract fields. Flask prototype using Sarvam AI.
Interruptible voice-agent runtime for structured interview prototypes, with VAD-based interruption handling and modular speech backends.
MCP Server for Brainiall Speech AI - pronunciation assessment, speech-to-text, and text-to-speech
Experimental Nepali TTS adaptation with Devanagari fixes, training tools, public model artifacts, and explicit evaluation limits.
Enterprise-Grade Secure ASR Diarization Pipeline - HIPAA-compliant speech processing service combining automatic speech recognition with speaker diarization. Features modular architecture, comprehensive security, and production-ready deployment.
Voice RAG that knows when not to answer - adaptive dense/sparse retrieval, real guardrails, and honest latency numbers. Sarvam (Hindi voice) + Gemini
Production-grade Speech AI platform with Whisper ASR, Pyannote Speaker Diarization, FastAPI, Streamlit, Docker, and multilingual speech processing.
To associate your repository with the speech-ai topic, visit your repo's landing page and select "manage topics."