Nemotron Labs VoiceChat is a full-duplex speech-to-speech model with realtime audio, transcripts, and tool calling. This guide converts the model, starts a local server, and connects a microphone client.
- An NVIDIA GPU with CUDA and at least 16 GB of GPU-accessible memory for the recommended Q4_K_M model
- The CUDA toolkit
- Git, Python 3.10+, CMake 3.26+, Ninja, and a C++17 compiler
On Ubuntu/Debian, install the remaining system packages with:
sudo apt-get update
sudo apt-get install -y \
build-essential cmake ninja-build git curl pkg-config libsentencepiece-dev \
libsndfile1 portaudio19-dev python3-venvSee Build from source for other platforms and toolchains.
git clone https://github.com/NVIDIA/NeMo-Speech.cpp.git
cd NeMo-Speech.cpp
git submodule update --init llama.cpp third_party/cpp-httplib
python3 -m venv .venv
. .venv/bin/activate
python -m pip install \
-r requirements.txt \
-r clients/voicechat/requirements.txtPyAudio is only needed for microphone capture and playback.
python convert_model.py \
nvidia/NVIDIA-NemotronLabs-VoiceChat-11B \
--outfile models/voicechatThis creates the recommended Q4_K_M bundle. Downloads use the Hugging Face
cache, so unchanged files are reused. If access is denied, accept the model
terms on Hugging Face and run hf auth login before trying again.
See Model conversion for NVFP4, BF16, local checkpoints, and other options.
scripts/configure.sh cuda-s2s
cmake --build --preset cuda-s2sStart a local server for one conversation:
build/cuda-s2s/bin/nemo-speech serve \
--s2s-model-dir models/voicechat \
--s2s-max-streams 1The server listens on localhost:8080. When model loading finishes, its
readiness endpoint is available at
http://localhost:8080/v1/realtime/health.
In another terminal:
cd NeMo-Speech.cpp
. .venv/bin/activate
python clients/voicechat/nemotron-voicechat-client.py \
--server ws://localhost:8080Press Ctrl+C to end the session. The client saves the response audio, transcripts, tool calls, and conversation log.
To use a WAV file without microphone or playback dependencies:
python clients/voicechat/nemotron-voicechat-client.py \
--server ws://localhost:8080 \
--input-file input.wav \
--audio-output response.wav \
--no-playback