This directory contains the OpenAPI 3.1 description for the built-in --server
mode.
s2-openapi.yaml- machine-readable API description forPOST /generate
The current server exposes a single endpoint:
POST /generate- synthesize speech from text, with optional reference audio for voice cloning
By default the response is a finalized audio/wav file. Real-time chunked
PCM16 WAV transport is used only when params.stream=true and
params.chunked=true (or params.realtime=true). That live transport starts
with a provisional WAV header because the server cannot seek back over an open
HTTP stream; saved files should still decode, but the header is not fully
finalized. params.output_format="pcm_s16le" skips the WAV container and
returns raw mono PCM16 bytes instead, which is easier for clients that keep a
dynamic playback buffer. The server handles one active synthesis at a time and
returns HTTP 503 when another request is already in progress.
Canonical multipart fields:
textreferencereference_textvoicevoice_dirparams
Accepted aliases:
reference:reference_audio,prompt_audio,ref_audioreference_text:ref_text,prompt_textvoice:voice_id,voice_profile
The params field is a JSON string. The OpenAPI schema documents the supported
keys currently parsed by the server:
max_new_tokenstemperaturetop_ptop_kmin_tokens_before_endn_threadscodec_follow_backendcodec_auto_backendcodec_decode_context_framescodec_context_frames(alias)voicevoice_idvoice_dirstreamchunkedrealtimestream_decode_stride_framesstream_decode_stridestream_holdback_framesstream_start_buffer_mssegment_sentencessentence_pause_mssegment_max_charslow_latencyoutput_formatverbose
Notes:
codec_follow_backend=trueasks the codec to follow the selected GPU backend when possible; if codec GPU initialization or allocation fails, the runtime falls back to CPU.codec_auto_backend=truelets the runtime benchmark/select the codec backend automatically. When disabled,codec_follow_backendis applied directly.voice=hoperesolves to./voices/hope.s2voiceby default.voice=voices/hope.s2voiceuses that exact file by inferringvoice_dirandvoice_idfrom the path.stream_decode_stride_frames=0keeps the server auto cadence, currently4frames.stream_decode_strideis an alias forstream_decode_stride_frames.stream_holdback_frames=0emits audio immediately after decode instead of waiting for the codec's full stability window. That is the lowest-latency option, but it can make chunk boundaries less stable.stream_start_buffer_ms=4000is a good starting point for natural playback withffplayor similar pipe readers: the server waits for about four seconds of PCM before it begins the chunked response.segment_sentences=trueswitches from exact frame streaming to sentence-by-sentence synthesis, which is usually the most natural option on machines that cannot keep the exact streaming path under real time.sentence_pause_ms=180adds a short fixed pause between segmented sentences.segment_max_charscan force very long sentences to be broken into smaller clauses before synthesis.low_latency=trueis a convenience preset for live playback: it defaults to stride1and holdback0unless you already overrode them explicitly.output_format="pcm_s16le"returns raw PCM16 bytes and addsX-Audio-Sample-Rate,X-Audio-Channels, andX-Audio-Encodingheaders so the client can configure playback without parsing a WAV header. Clients should use those headers instead of assuming a fixed sample rate.
curl -X POST http://127.0.0.1:3030/generate \
--form "text=Hello world" \
--form 'params={"temperature":0.9,"top_p":0.9}' \
-o output.wavReference audio must be accompanied by its transcript via reference_text or
one of its accepted aliases. Request validation failures, JSON parse failures,
and invalid params value types return 400; synthesis failures detected
before the response body starts return 500. Real-time chunked responses may
instead terminate early if generation fails after streaming has already begun.