WaveKat’s cover photo
WaveKat

WaveKat

Software Development

Auckland, Auckland 3 followers

The Conversational AI Platform for Task Automation.

About us

WaveKat is an open-source voice AI platform built in Rust, providing the foundational building blocks for intelligent phone and voice communication systems.

Website
https://wavekat.com
Industry
Software Development
Company size
1 employee
Headquarters
Auckland, Auckland
Type
Self-Employed

Locations

Employees at WaveKat

Updates

  • WaveKat Voice can now be driven by an AI assistant. You can ask something like Claude to "call the dentist and wait until someone picks up," and it dials through the app, follows the call, and tells you how it went. The assistant handles the dialing and waiting — you still do the talking. A bit more on what it does: - Place and manage real phone calls from a command-line tool or MCP server - Wait for an outcome, list active calls, send touch-tones to navigate phone menus, pull the transcript, hang up - JSON output and clean exit codes, so an assistant can branch on whether a call was answered, busy, failed, or unanswered - One-click connect for Claude Desktop, Claude Code, Cursor, Codex, Gemini, and Windsurf - Nothing extra to install — the tool ships inside the app; off by default and opt-in Runs on Mac and Linux today. Wrote up how it works here: https://lnkd.in/ex_gJDju #VoiceAI #AIAgents #MCP #PhoneAutomation #SIP

  • The best part of building in public? The first user request. 🚀 We recently reached a major milestone for wavekat-tts (https://lnkd.in/ezEuZnsM). A user opened our first-ever feature request, asking for Qwen3-TTS 0.6B Base support to enable high-fidelity voice cloning. We believe in moving fast. We got to work immediately and are excited to announce that the feature is already live! WHAT IS WAVEKAT-TTS? It is a high-performance, Rust-based Text-to-Speech (TTS) library designed for speed and ease of use. It allows developers to integrate advanced speech synthesis and voice cloning into their applications with minimal configuration. NEW IN v0.0.3 (SHIPPED TODAY): - VOICE CLONING: Full integration of the Qwen3-TTS 0.6B Base model. - PRECISION CONTROL: Support for INT4 precision for blazing speed and FP32 for maximum fidelity. - AUTOMATED WORKFLOW: Built-in auto-resampling for reference audio and automatic model downloading. Closing our first community "Issue" on the same day it was opened is a proud moment for us. A huge thank you to our early users for pushing wavekat forward. Check out the update and the code example here: https://lnkd.in/eY84X6ZT #BuildInPublic #OpenSource #AI #RustLang #VoiceCloning #MachineLearning #TTS

  • When should a voice agent start talking? In conversational AI, turn detection is one of the hardest unsolved UX problems. Respond too early → you interrupt the user. Too late → awkward silence. Pipecat Smart Turn is one of the leading approaches to solving this. But how well does it actually work in practice? I built WaveKat Lab to answer exactly this kind of question — visually. In this video, I use WaveKat Lab to test Pipecat Smart Turn v3 with: → Live recording analysis → Per-frame prediction visualization with confidence scores & inference latency → VAD-Gated Pipeline mode simulating real production workflows The goal: make speech model behavior observable and debuggable, not a black box. WaveKat is an open-source speech processing toolkit I'm building in Rust. WaveKat Lab is its browser-based visual experimentation tool. 🎥 Watch the full test: https://lnkd.in/eMp9sNKV 🌐 Website: https://wavekat.com 💻 GitHub: https://lnkd.in/eHseddz8 If you're building voice agents or working on conversational AI, I'd love to hear how you're tackling turn detection in your systems. #PipecatSmartTurn #WaveKat #VAD #TurnDetection #VoiceAgent #SpeechAI #ConversationalAI #OpenSource

  • WaveKat.com is live. We're building open-source, AI-powered tools that give every small business the voice of a big one. Voice is where we start — answering phones, handling conversations, being present 24/7. Capabilities that used to require enterprise budgets, now open to everyone. Our Rust-based libraries are already on GitHub:   - wavekat-core — audio primitives  - wavekat-vad — voice activity detection  - wavekat-turn — turn detection  - wavekat-lab — interactive testing dashboard  All Apache 2.0. Check it out at https://wavekat.com/  #opensource #ai #voiceai #rust #smallbusiness

  • View organization page for WaveKat

    3 followers

    Introducing wavekat-vad — a unified Voice Activity Detection library for Rust 🦀 We just published wavekat-vad on crates.io. VAD (Voice Activity Detection) is the critical first step in any voice AI pipeline — determining when someone is speaking vs. silence. Get it wrong, and everything downstream (ASR, LLM, TTS) suffers. But the VAD landscape is fragmented: different models, different APIs, different trade-offs between speed and accuracy. wavekat-vad solves this with one unified Rust API across multiple backends. Pick the model that fits your use case with a Cargo feature flag. No code changes to switch. Currently supported backends: → WebRTC VAD — GMM-based, ultra-low latency, near-zero overhead → Silero VAD — neural network, widely adopted, strong accuracy → TEN VAD — enterprise-grade DNN, optimized for low-latency turn detection → FireRedVAD — the newest addition. 97.57% F1 on FLEURS-VAD-102, 100+ languages, detects speech, singing, and music. From Xiaohongshu's Super Intelligence Lab. And more to come. Get started in one line: > cargo add wavekat-vad --features silero We built this because we needed it ourselves — wavekat-vad is a core component of WaveKat, a voice AI platform we're building in Auckland, New Zealand. We decided to open-source it as a standalone crate so anyone building real-time voice pipelines in Rust can benefit. Open source, Apache-2.0. 📦 https://lnkd.in/e2vavAUy 📖 https://lnkd.in/eq2HpfUS 🔗 https://lnkd.in/eeC9xJkp #Rust #VoiceAI #VAD #OpenSource #WaveKat #SpeechProcessing

    • No alternative text description for this image

Similar pages