Skip to content

Latest commit

 

History

21 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

NeMo-Speech.cpp

A lightweight native C++ runtime for the NVIDIA Nemotron Speech model family, with broad hardware support.

Models  |  Installation  |  Quick Start  |  Documentation  |  API Reference

NeMo-Speech.cpp is NVIDIA's official solution for local speech inference, providing day-0 support for the latest models from NVIDIA NeMo Speech, with native inference powered by ggml.

Models and applications

Application Supported models
Speech recognition Nemotron 3.5 ASR Streaming 0.6B, Nemotron Speech Streaming 0.6B, Parakeet TDT 0.6B v3, and Parakeet CTC 1.1B
Speaker diarization Streaming Sortformer 4-speaker v2 and Nemotron 3 Diarization, standalone or combined with ASR
Text and speech translation Riva Translate 4B Instruct v2, with composed ASR-to-NMT-to-TTS speech translation
Speech synthesis MagpieTTS Multilingual 357M with NeMo NanoCodec
Full-duplex voicechat Nemotron Labs VoiceChat, including realtime audio, transcripts, and tool calling
Speech processing Silero VAD, punctuation and capitalization, endpointing, text normalization, and subtitles

Performance

NeMo-Speech.cpp is blazing fast and built for real-time streaming. Speech recognition transcribes audio in 160 ms chunks up to 67× faster than real time, and speech synthesis generates speech up to 60× faster than real time with its first audio in under 10 ms. Both stay faster than real time even on a CPU.

Streaming speech recognition with Nemotron Speech Streaming 0.6B (Q8_0), 160 ms chunks:

Device Latency per chunk Throughput Speedup over NeMo (FP32)
GeForce RTX 4090 2.3 ms 67× real time 8.8×
CPU 27 ms 6× real time 3.7×

Streaming speech synthesis with MagpieTTS Multilingual (Q8_0), 186 ms audio chunks:

Device Time to first audio Inter-chunk latency Throughput Speedup over NeMo (FP32)
GeForce RTX 4090 9 ms 3 ms 60× real time 40×
CPU 203 ms 64 ms 2.7× real time 9.7×

See BENCHMARK.md for the methodology and more results.

Installation

Important

For the best performance, build natively from source. A native build is compiled for your machine. Tagged releases are cut periodically and can trail the main branch; for prebuilt binaries of the latest main, pass --channel nightly (-Channel nightly on Windows) to install the nightly prerelease, which is rebuilt daily. See Build from source.

Install the nemo-speech CLI for the detected platform and backend:

On Linux or macOS, run:

curl -fsSL https://github.com/NVIDIA/NeMo-Speech.cpp/raw/main/scripts/install.sh | sh

Open a new shell after installation so the updated user PATH takes effect.

On Windows, run from PowerShell:

irm https://github.com/NVIDIA/NeMo-Speech.cpp/raw/main/scripts/install.ps1 | iex

Open a new PowerShell window after installation so the updated user PATH takes effect.

The installer downloads the prebuilt archive for the latest release, checks it against the SHA-256 checksum published with the release, and builds from source when no archive is available for your platform. Pass --source (-Source on Windows) to always build from the main branch. A source build requires Git, CMake 3.26 or newer, Ninja, a C++17 compiler, SentencePiece development files, and the toolchain required by the selected backend, if any. See Installation for platform-specific prerequisites and options.

Quick start

Transcribe a local WAV file. On first use, the CLI downloads the pinned default Nemotron 3.5 GGUF from Hugging Face and verifies its size and SHA-256:

nemo-speech transcribe /path/to/audio.wav

Source checkouts can use test_files/asr/wav/test/jfk.wav as a smoke-test input.

The same command can transcribe the default microphone on builds that include live capture:

nemo-speech transcribe --live

Run nemo-speech model list to see defaults, short names, and which command uses each model. For example, nemo-speech pull nemotron-en downloads the English-only model ahead of time, and --model nemotron-en selects it. Local GGUF paths continue to work without downloading anything. The CLI selects an available backend and handles common mono or stereo PCM WAV sample rates automatically. See the CLI model guide and model conversion for custom checkpoints.

Command line

The CLI is the primary interface. Run nemo-speech --help to see the capabilities included in your build. The CLI guide covers model selection, GPU controls, directory transcription, subtitles, diarization, translation, synthesis, structured output, and benchmarking. The VoiceChat guide covers building and serving the realtime pipeline.

Local server and playground

Start the same runtime as a local HTTP service and open the playground:

nemo-speech serve \
  --asr-model nemotron-3.5 \
  --open

The server binds to http://127.0.0.1:8080 by default. Its transcription and speech routes expose documented OpenAI-compatible subsets, alongside realtime WebSocket transcription and realtime VoiceChat when its model is loaded. A separately built riva_server binary provides the ASR, TTS, and translation gRPC interfaces. See the server guide when you are ready to integrate either frontend.

Native SDK

Release archives include stable C headers, shared libraries, and an exported CMake package. An installed application can link only the capability it uses:

find_package(NeMoSpeech REQUIRED COMPONENTS ASR)
target_link_libraries(my_app PRIVATE NeMoSpeech::ASR)

See native SDK integration for in-process C/C++ usage, or client integration for OpenAI SDK, curl, and Riva-compatible gRPC usage.

Build from source

Requires CMake 3.26 or newer, Ninja, C and C++17 compilers, SentencePiece development files, and the toolchain required by the selected backend, if any. For a CUDA ASR and TTS server with the playground:

git submodule update --init llama.cpp third_party/cpp-httplib
scripts/configure.sh cuda-server
cmake --build --preset cuda-server

The configuration helper validates required submodules and applies the pinned ggml patch series for CUDA builds. CPU, Metal, Vulkan, server, component, Windows, and container instructions are in Build from source.

Documentation

Start here What it covers
Installation Native releases, Windows, upgrades, and manual verification
CLI guide Transcription, subtitles, directories, diarization, NMT, TTS, and tooling
Model conversion Convert NeMo and Hugging Face checkpoints to runtime GGUF files
Speech-to-speech VoiceChat Convert, serve, and exercise the streaming S2S pipeline
Servers HTTP playground/realtime serving and the separate Riva-compatible gRPC server
HTTP API reference Every endpoint's request fields, responses, and the realtime protocol
Native SDK CMake components, C ABI lifetimes, threading, and examples
Client integration OpenAI SDKs, curl, and Riva gRPC clients
Troubleshooting doctor output and common runtime failures
Build from source Presets, optional components, dependencies, containers, and artifacts
All documentation ASR, TTS, NMT, configuration, and developer references

License

NVIDIA-authored code is released under the Apache License 2.0, with the project copyright notice in NOTICE. Third-party components retain their respective terms; see Third-Party Notices. Release archives also include these files under share/licenses/nemo-speech/.

Contributing

External contributions are welcome. See CONTRIBUTING.md for the contribution terms and Developer Certificate of Origin sign-off process.

About

NeMo-Speech.cpp is a lightweight C++ inference runtime for Speech models

Topics

Resources

Contributing

Security policy

Stars

154 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages