AudioShake builds AI models for separating and understanding sound.
Our technology turns mixed audio into usable components and structured data for music, media, speech, and machine-learning workflows.
Separate multi-speaker recordings into individual speaker tracks, including overlapping speech, with diarization and confidence scores.
- Multi-Speaker 2.0 Technical Evaluation
- Multi-Speaker Samples on Hugging Face
- Multi-Speaker API documentation
Isolate dialogue and speech from background noise, music, and other interference for transcription, voice AI, media, and real-time applications.
AudioShake provides tools for separating, detecting, and identifying audio in film, television, broadcast, and other media workflows.
- Remove Dialogue for Dubbing
- Remove Music from Content for Copyright Compliance
- Detect Music in Content
- Identify Music
Separate, transcribe, and transform music for production, interactive, and creator workflows.
We build and contribute to open evaluation tools and benchmarks for audio AI.
ALT-Eval is an evaluation toolkit for Automatic Lyrics Transcription (ALT), designed to measure both transcription accuracy and readability.
JAM-ALT is a community benchmark for automatic lyrics transcription, developed by AudioShake and Spotify.
It provides human-transcribed lyrics with word-level timestamps for evaluating lyrics transcription and alignment systems.