Confucius4-R2T2 (Real Real-Time Transcription) is NetEase Youdao's low-latency, append-only streaming ASR model. It is a Qwen3-ASR-1.7B fine-tune that keeps the same audio tower and thinker graph, and adds a Longest Stable Prefix (LSP) decoding algorithm: every chunk re-decodes the whole accumulated audio with the previously recognized text as a continuation prompt, and only the stable prefix of the new result is committed downstream. Committed text is never revised, which is what makes it suitable for live captioning and downstream agents.
| Field | Value |
|---|---|
| Family | confucius4_r2t2 |
| HF checkpoint | netease-youdao/Confucius4-R2T2 |
| Task | asr |
| Modes | offline, streaming |
| Input | 16 kHz speech WAV (other rates are resampled by the frontend) |
| Output | Transcript text |
| Streaming output | Append-only partial text plus the final transcript |
| Timestamps | Not supported |
| Context / hotwords | Optional --text system prompt |
confucius4_r2t2 is a standalone family (a community port, status: community in
the spec): assets, Whisper log-mel frontend, windowed audio tower, tokenizer, LSP
session, and text post-processing live under
src/community_models/confucius4_r2t2/ + include/engine/community_models/confucius4_r2t2/ and
share no code with the qwen3_asr family, so the two can evolve independently.
The loader itself is the framework's schema-v1 spec-backed loader —
model_specs/confucius4_r2t2.json is the single source of truth for metadata,
capabilities, options and packages, and the factory lives next to the session
(make_confucius4_r2t2_loader), so there is no per-model loader.{h,cpp}. What it does
not duplicate is framework infrastructure:
- the thinker is a thin adapter over the shared greedy Qwen decoder runtime
(
runtime::GreedyQwenDecoderRuntime), which owns the prefill/decode graphs, the audio-embedding injection, the static KV cache, and greedy sampling; - the audio tower builds on the shared modules (
Conv2dModule,GeluModule,LinearModule,LayerNormModule,ScaledDotProductAttentionModule) and the shared layout/contiguity helper (core::ensure_backend_addressable_layout); - long-audio chunking uses the shared audio chunk planner
(
engine::audio::plan_audio_chunks+slice_audio_buffer); - multichannel and off-rate input goes through the shared mono conversion and resampling helper inside the frontend.
The family-specific pieces are:
session.cpp— offline transcription plus the LSP streaming state machine,text_postprocess.cpp—normalize_punct_by_context, repetition repair,language X<asr_text>parsing, Chinese spacing, and|truncation,assets.cpp— checkpoint resolution that keeps symlinked weight files loadable and rejects precisions below Q8_0 (see below),thinker.cpp— maps the R2T2 config andthinker.*tensor layout onto the shared decoder, including the tied LM head the checkpoint ships without.
The R2T2 config declares interleaved mrope with mrope_section [24, 20, 20], but
the model only ever sees audio and text, so all three position streams are
identical and mrope degenerates to standard NEOX RoPE. The shared decoder's
NEOX rope therefore reproduces the reference numerics, which the golden checks
in tests/confucius4_r2t2/ confirm chunk by chunk.
python3 tools/model_manager_v2.py install confucius4_r2t2_safetensorsOr point --model at any directory with the HF checkpoint layout (config.json,
generation_config.json, preprocessor_config.json, model.safetensors,
tokenizer files). Hugging Face cache snapshot directories work directly, even
though their model.safetensors is a symlink into the blob store: the family
resolves the checkpoint itself instead of going through the shared canonicalized
tensor path.
audiocpp_cli --task asr --family confucius4_r2t2 \
--model models/Confucius4-R2T2 --backend metal \
--audio speech_16k.wav --text-out transcript.txtAudio longer than 30 seconds is split into 30-second chunks inside the session
and the chunk transcripts are joined with a single space. Pass --language Chinese (or any supported language) to skip language detection, and --text "hotword, another term" to bias recognition with a context prompt.
audiocpp_cli --task asr --mode streaming --family confucius4_r2t2 \
--model models/Confucius4-R2T2 --backend metal \
--audio speech_16k.wav \
--session-option confucius4_r2t2.chunk_size_ms=320 \
--text-out transcript.txtIn streaming mode committed text appears as it stabilizes and the complete
transcript is returned when the stream ends. --audio - streams raw 16 kHz mono
PCM from stdin for live sources.
Chunk sizes from 80 ms to 2000 ms are supported. 320 ms is a good default on Apple Silicon: 160 ms trades accuracy for latency and runs below real time with eager Metal execution on smaller machines, while 2000 ms approaches offline quality.
| Option | Values | Default | Meaning |
|---|---|---|---|
confucius4_r2t2.chunk_size_ms |
80-2000 | 320 |
Streaming decode chunk in milliseconds. |
confucius4_r2t2.unfixed_chunk_num |
integer | 2 |
Leading chunks decoded without a stable-prefix prompt. |
confucius4_r2t2.unfixed_token_num |
integer | 5 |
Tokens rolled back from the accumulated text before it is used as the prefix prompt. |
confucius4_r2t2.rollback_punctuation |
true, false |
false |
Keep trailing text uncommitted when it already ends with punctuation instead of rolling back tokens. |
confucius4_r2t2.max_tokens |
integer | 32 |
Greedy decode budget per chunk and for the final flush. |
confucius4_r2t2.audio_encoder_weight_type |
native, f32, f16 |
native |
Audio tower weight storage. |
confucius4_r2t2.thinker_weight_type (alias confucius4_r2t2.weight_type) |
native, f32, f16, bf16, q8_0 |
native |
Thinker weight storage. |
confucius4_r2t2.audio_encoder_graph_arena_mb |
MB | 128 |
Audio tower graph arena. |
confucius4_r2t2.thinker_prefill_graph_arena_mb |
MB | 256 |
Thinker prefill graph arena. |
confucius4_r2t2.thinker_decode_graph_arena_mb |
MB | 256 |
Thinker decode graph arena. |
confucius4_r2t2.thinker_weight_context_mb |
MB | 64 |
Thinker weight context. |
| Option | Values | Default | Meaning |
|---|---|---|---|
language (or --language) |
language name | empty | Force the recognition language; Auto keeps detection. |
max_tokens (or --max-tokens) |
integer | model config | Offline decode budget. |
Each full chunk:
- Appends the chunk to the accumulated audio (nothing is dropped, no padding).
- Builds the prompt as
chat template + stable prefix, where the stable prefix is the accumulated transcript minusunfixed_token_numtokens. Beforeunfixed_chunk_numchunks the prefix is empty. Decoding a prefix never cuts a multi-byte token: the rollback grows until the decoded prefix contains no U+FFFD replacement character. - Greedy-decodes with a
max_new_tokensbudget on the accumulated audio. - Normalizes punctuation by context, reparses the
language X<asr_text>textoutput, and re-derives the stable prefix (again minusunfixed_token_numtokens). - Reports the growth of the stable prefix as committed text.
finish_stream flushes any tail shorter than one chunk with a fixed rollback,
then returns the complete transcript.
Re-decoding the whole accumulated audio is the reference algorithm, but not the reference cost. Two properties keep each chunk bounded:
- The audio tower is block-diagonal per
n_window_inferattention window (4 s), with per-chunk convolutions and positions, so embeddings of completed windows never change. They are cached, and only the trailing partial window is re-encoded (together with the previous window when the tail is shorter than one conv chunk, so the reusable graph is kept). - The thinker prompt is
template head, audio tokens, template close, stable prefix. The K/V rows for the head and the cached audio tokens are retained in the decode cache between chunks (GreedyQwenDecoderRuntime::generate_incremental), so only the tail audio, the close, and the stable text prefix are prefilled. The cache grows in 512-step buckets; crossing one falls back to a full prefill for that chunk.
Per-chunk cost is therefore dominated by the stable text prefix and the greedy tail decode, not by the utterance length.
Peak normalization and the log-mel floor are computed over all accumulated audio, so a louder segment rewrites the features of earlier windows. Cached windows are therefore reused only while their log-mel features are bit-identical to the current ones; otherwise the encoder cache and the thinker rows of those audio tokens are dropped and rebuilt, and that chunk costs a full re-encode and prefill.
Committed deltas contain transcript text only. In automatic language mode,
a rolled-back prefix is not publishable until it contains <asr_text>;
metadata fragments such as language neither emit a delta nor advance the
published offset. Forced-language prompts already supply the metadata, so
untagged decoded prefixes remain valid transcript text in that mode.
After metadata removal, growing stable prefixes emit only the suffix beyond
the previously published length, counted in Unicode code points. A shrinking
prefix emits nothing. The uncommitted tail is delivered in the final result
(transcript.text.done on the server), rather than as a final delta.
The reference also ships a rolling-window variant for unbounded streams ("no reset": keep 16 s of audio, discard the oldest 8 s and the matching text). This port implements the standard variant used by the upstream WebSocket server, which bounds audio per utterance with VAD. Because the audio tower uses 1500 positions, a single unsegmented stream is limited to roughly 110 s of accumulated audio; segment longer streams (as the reference server does) or add the rolling-window variant.
Streaming no longer rebuilds the major inference graphs on every chunk.
Opt-in via encode(features, reuse_graph) and
R2T2ASRGenerationOptions::reuse_graphs (the streaming session enables both;
the offline path keeps exact per-frame graphs):
- The audio encoder builds one graph sized to a capacity bucket (one to two chunks exactly, then four-chunk steps) and refills the attention mask for the valid token prefix on each run, so a growing stream reuses the graph until the next bucket boundary.
- The thinker prefill runs through the shared Qwen chunked prefill runtime
(64-token blocks, one reserved graph) and the decode graph keeps a KV-cache
capacity grown in 128-token buckets up to
max_position_embeddings; new prompts clear KV on device. - Prompt token embeddings reuse a fixed-width lookup graph instead of a per-request build.
The Metal encoder avoids padding an input with fewer than 64 tokens to a capacity of 64 tokens or more, and rebuilds on shrink across that boundary. The bundled Metal backend switches attention value multiplication to a half-input SIMD-group matrix kernel at that size. A larger bucket can therefore change arithmetic precision, not just floating-point summation order.
Diagnostics on the M3 found identical convolution output, first-layer norm, and Q/K/V for the tested inputs; the first difference appeared in attention output when this boundary was crossed. Padding is not inherently inexact: some padded shapes matched bit for bit. These findings describe the tested backend and shapes, not a proof that all padding differences have one cause.
test_confucius4_r2t2_graph_reuse reads the actual retained graph capacity,
including after shrink. It checks the Metal precision boundary, unpadded
output equality, padded repeatability, and a backend-specific
relative-RMSE drift alarm on the synthetic fixture (2e-3 on Metal under
the 64-token guard, 2.5e-2 on CPU, where fp32 reduction-order noise is
larger). That alarm is not a transcript-quality guarantee.
Joint tests pass exact encoder output through the exact decoder and reused
encoder output through the reused decoder, requiring identical token IDs for
real English audio prefixes with automatic and forced-English prompts. They
also check nonempty output, bucket boundaries, and growth followed by shrink.
The separate decoder-only test continues to cover injected prompt lengths.
For a repeatable local benchmark, use the same checkpoint, backend, thread
count, and audio on both revisions. Count positive graph.build_ms entries
in --log-file output and compare confucius4_r2t2.session.stream.wall_ms,
CLI deltas, final text, and process memory measurements. The bundled English
sample is 14.0719 seconds (44 decode calls including the final flush at the
default 320 ms chunk size). A single timing comparison is not a general
performance guarantee; peak memory footprint and maximum resident set size
are different measurements.
The port still recomputes the accumulated audio and prompt on each chunk; incremental encoder/prefill caching is a separate task.
When the audio ends right after the last word, the model often drops that
word. Push-to-talk dictation hits this when the key is released as the word
ends. Both modes are affected, so it is not a streaming chunk-boundary effect.
With v0.9.0, Metal, the Q8_0 GGUF and --language English, on short
16 kHz dictations:
| Audio | Offline | Streaming (320 ms) |
|---|---|---|
| As recorded (0.9 s) | It seems |
It seems |
| Plus 300 ms of silence | It seems fixed. |
It seems fixed. |
On the live server route, another 1.1 s clip gave Thank you very and
Thank you very much. with 200 ms or more of silence. This has not been
compared against the reference implementation yet.
Clients that end a stream when the user stops speaking should append about
300 ms of zero samples before closing the request body (or before EOF on
--audio -).
server.local.json declares r2t2-asr (offline) and r2t2-asr-stream
(streaming, 320 ms chunks):
curl http://127.0.0.1:8488/v1/audio/transcriptions \
-F model=r2t2-asr -F file=@speech.wav# Decoding deltas of an uploaded file
curl -N http://127.0.0.1:8488/v1/audio/transcriptions \
-F model=r2t2-asr-stream -F stream=true -F file=@speech.wav# Live PCM: deltas appear while the audio is still arriving
ffmpeg -f avfoundation -i ":0" -ar 16000 -ac 1 -f s16le - \
| curl -N -X POST -H 'Expect:' -T - \
'http://127.0.0.1:8488/v1/audio/transcriptions/live?model=r2t2-asr-stream&sample_rate=16000&channels=1&sample_format=s16le'Add &prompt=<url-encoded hotwords> to bias the live stream the same way --text does on the CLI.
GGUF is supported for f32, f16, bf16, q8_0 and the k-quants q2_k,
q4_k and q6_k. f16 and q8_0 reproduce the reference transcripts
exactly; storage types outside that set (legacy q4_0, q5_k, ...) are
rejected at load time with an actionable error instead of silently decoding to
empty text, because this graph's kernels are not validated for them.
Verified conversions are published at
davidxifeng/Confucius4-R2T2-gguf
(a quantized derivative work, distributed under the upstream NetEase Youdao
model license — see the repository's LICENSE, LICENSE_zh, and NOTICE):
| Package id | File | Quantization |
|---|---|---|
confucius4_r2t2_q8_0 (default) |
r2t2-q8_0.gguf |
Q8_0 |
confucius4_r2t2_f16 |
r2t2-f16.gguf |
F16 |
For current file sizes, SHA-256 checksums, runtime compatibility, and published checkpoint validation, see the Hugging Face model card. Release-specific artifact details are maintained there rather than duplicated in this document.
A community Q4_K_M quantization by NairoDorian
is published at
Nairod785/Confucius4-R2T2-Q4_K_M-GGUF
and pinned to a fixed revision. It is a quantized derivative work under the
same NetEase Youdao model license (LICENSE, LICENSE_zh, NOTICE in that
repository); QUANTIZATION.md there documents the per-tensor recipe and the
measured speed/quality results.
| Package id | File | Quantization |
|---|---|---|
confucius4_r2t2_q4_k_m |
r2t2-q4_k_m.gguf |
Q4_K_M (Q4_K with Q6_K MLP-down and Q2_K embeddings; BF16 audio tower) |
python3 tools/model_manager_v2.py install confucius4_r2t2_q4_k_mpython3 tools/model_manager_v2.py install confucius4_r2t2_q8_0 # or confucius4_r2t2_f16These files are self-contained: tokenizer, processor/generation config, chat
template and the model spec are embedded, and the embedded spec carries the
schema-v1 option contract, so option validation comes from the file itself
rather than from a spec installed beside the runtime (a legacy embedded spec
triggers a [warning][model_spec] fallback instead).
Convert a real (non-symlinked) checkpoint directory:
audiocpp_gguf \
--input /path/to/Confucius4-R2T2/model.safetensors \
--root /path/to/Confucius4-R2T2 \
--family confucius4_r2t2 \
--output models/Confucius4-R2T2-GGUF/r2t2-q8_0.gguf \
--type q8_0The converter resolves sidecars from a real directory, so run it against a
directory whose files are not symlinks (a plain cp -r of the checkpoint, or a
directory created by the model manager). Direct loading of a symlinked HF cache
snapshot is supported and does not need this step. The output embeds the
sidecars and the model spec, so the single .gguf is portable:
audiocpp_cli --task asr --family confucius4_r2t2 \
--model models/Confucius4-R2T2-GGUF/r2t2-q8_0.gguf --backend metal \
--audio speech_16k.wavLoading accepts either the .gguf file, a directory holding one, or a directory
holding both formats — a GGUF checkpoint wins when both are present, matching
the Qwen-family convention.
The published checkpoint ties the LM head to the token embedding and therefore
contains no lm_head.weight; the family detects that and reuses the embedding,
so conversions need no special flags.
The port is verified against the macOS MPS reference with golden traces:
# 1. Golden from the reference implementation (in the Confucius4-R2T2 repo)
cd /path/to/Confucius4-R2T2
PYTHONPATH=. uv run python /path/to/audio.cpp/tests/confucius4_r2t2/make_golden.py \
--model_path /path/to/audio.cpp/models/Confucius4-R2T2 \
--audio resources/test.wav \
--out /path/to/audio.cpp/tests/confucius4_r2t2/golden.json
# 2. Compare the C++ runtime (offline text, per-chunk committed text, final text)
cd /path/to/audio.cpp
python3 tests/confucius4_r2t2/compare.py \
--cli build/macos-metal-release/bin/audiocpp_cli \
--model models/Confucius4-R2T2 \
--audio /tmp/test.wav \
--golden tests/confucius4_r2t2/golden.json --backend metalcompare.py runs the CLI with trace logging, parses the per-chunk trace, and
diffs it against the golden. A repo-native smoke test runs both paths without
the Python environment:
build/macos-metal-release/bin/test_confucius4_r2t2_transcription --backend metal(It skips with exit code 125 when models/Confucius4-R2T2 or the audio asset is
absent.)
compare.py checks four things per golden: the offline transcript, the
committed delta stream, the final streaming transcript, and every per-chunk
fixed_text.
| Golden | Offline text | Committed stream | Final transcript | Per-chunk fixed_text |
|---|---|---|---|---|
golden.json (upstream Chinese resources/test.wav) |
exact | exact | exact | 21/21 exact |
golden_zh.json (same audio, --language Chinese) |
exact | exact | exact | 21/21 exact |
golden_sample16k.json (assets/resources/sample_16k.wav, English) |
exact | exact | exact | 41/43 exact |
golden_sample16k_bf16.json (same audio, bf16 reference) |
exact | exact | exact | 41/43 exact |
These are the original port's reference-parity results. The regression
comparison now filters metadata-only prefixes from the unmodified Python
goldens before checking committed deltas and per-chunk fixed_text. The C++
stream intentionally excludes the reference's language artifact; offline
and final transcript expectations are unchanged.
On the 14 s English clip, chunks 37 and 42 differ in fixed_text by about one
token of the trailing numeric span (22,00 vs 22,0, and a trailing
you. predicted one chunk earlier). This is greedy-decoding sensitivity, not
port logic:
- both differences appear in the raw decoded text, before any R2T2 post-processing or rollback runs;
- the reference disagrees with itself at the same order of magnitude — its fp16 and bf16 runs differ at chunks 13 and 37;
- the difference never reaches the wire, because
finish_streampublishes no delta and the final transcript is identical.
So per-chunk fixed_text equality is exact for the Chinese reference clip and
stable to within one token for long English audio, with metadata artifacts excluded from committed output.