Skip to content

Latest commit

 

History

History
450 lines (362 loc) · 21.3 KB

File metadata and controls

450 lines (362 loc) · 21.3 KB

Confucius4-R2T2

Confucius4-R2T2 (Real Real-Time Transcription) is NetEase Youdao's low-latency, append-only streaming ASR model. It is a Qwen3-ASR-1.7B fine-tune that keeps the same audio tower and thinker graph, and adds a Longest Stable Prefix (LSP) decoding algorithm: every chunk re-decodes the whole accumulated audio with the previously recognized text as a continuation prompt, and only the stable prefix of the new result is committed downstream. Committed text is never revised, which is what makes it suitable for live captioning and downstream agents.

Field Value
Family confucius4_r2t2
HF checkpoint netease-youdao/Confucius4-R2T2
Task asr
Modes offline, streaming
Input 16 kHz speech WAV (other rates are resampled by the frontend)
Output Transcript text
Streaming output Append-only partial text plus the final transcript
Timestamps Not supported
Context / hotwords Optional --text system prompt

Independent implementation

confucius4_r2t2 is a standalone family (a community port, status: community in the spec): assets, Whisper log-mel frontend, windowed audio tower, tokenizer, LSP session, and text post-processing live under src/community_models/confucius4_r2t2/ + include/engine/community_models/confucius4_r2t2/ and share no code with the qwen3_asr family, so the two can evolve independently. The loader itself is the framework's schema-v1 spec-backed loader — model_specs/confucius4_r2t2.json is the single source of truth for metadata, capabilities, options and packages, and the factory lives next to the session (make_confucius4_r2t2_loader), so there is no per-model loader.{h,cpp}. What it does not duplicate is framework infrastructure:

  • the thinker is a thin adapter over the shared greedy Qwen decoder runtime (runtime::GreedyQwenDecoderRuntime), which owns the prefill/decode graphs, the audio-embedding injection, the static KV cache, and greedy sampling;
  • the audio tower builds on the shared modules (Conv2dModule, GeluModule, LinearModule, LayerNormModule, ScaledDotProductAttentionModule) and the shared layout/contiguity helper (core::ensure_backend_addressable_layout);
  • long-audio chunking uses the shared audio chunk planner (engine::audio::plan_audio_chunks + slice_audio_buffer);
  • multichannel and off-rate input goes through the shared mono conversion and resampling helper inside the frontend.

The family-specific pieces are:

  • session.cpp — offline transcription plus the LSP streaming state machine,
  • text_postprocess.cpp — normalize_punct_by_context, repetition repair, language X<asr_text> parsing, Chinese spacing, and | truncation,
  • assets.cpp — checkpoint resolution that keeps symlinked weight files loadable and rejects precisions below Q8_0 (see below),
  • thinker.cpp — maps the R2T2 config and thinker.* tensor layout onto the shared decoder, including the tied LM head the checkpoint ships without.

Rope note

The R2T2 config declares interleaved mrope with mrope_section [24, 20, 20], but the model only ever sees audio and text, so all three position streams are identical and mrope degenerates to standard NEOX RoPE. The shared decoder's NEOX rope therefore reproduces the reference numerics, which the golden checks in tests/confucius4_r2t2/ confirm chunk by chunk.

Install

python3 tools/model_manager_v2.py install confucius4_r2t2_safetensors

Or point --model at any directory with the HF checkpoint layout (config.json, generation_config.json, preprocessor_config.json, model.safetensors, tokenizer files). Hugging Face cache snapshot directories work directly, even though their model.safetensors is a symlink into the blob store: the family resolves the checkpoint itself instead of going through the shared canonicalized tensor path.

Offline transcription

audiocpp_cli --task asr --family confucius4_r2t2 \
  --model models/Confucius4-R2T2 --backend metal \
  --audio speech_16k.wav --text-out transcript.txt

Audio longer than 30 seconds is split into 30-second chunks inside the session and the chunk transcripts are joined with a single space. Pass --language Chinese (or any supported language) to skip language detection, and --text "hotword, another term" to bias recognition with a context prompt.

Streaming transcription

audiocpp_cli --task asr --mode streaming --family confucius4_r2t2 \
  --model models/Confucius4-R2T2 --backend metal \
  --audio speech_16k.wav \
  --session-option confucius4_r2t2.chunk_size_ms=320 \
  --text-out transcript.txt

In streaming mode committed text appears as it stabilizes and the complete transcript is returned when the stream ends. --audio - streams raw 16 kHz mono PCM from stdin for live sources.

Chunk sizes from 80 ms to 2000 ms are supported. 320 ms is a good default on Apple Silicon: 160 ms trades accuracy for latency and runs below real time with eager Metal execution on smaller machines, while 2000 ms approaches offline quality.

Session options (use with --session-option)

Option Values Default Meaning
confucius4_r2t2.chunk_size_ms 80-2000 320 Streaming decode chunk in milliseconds.
confucius4_r2t2.unfixed_chunk_num integer 2 Leading chunks decoded without a stable-prefix prompt.
confucius4_r2t2.unfixed_token_num integer 5 Tokens rolled back from the accumulated text before it is used as the prefix prompt.
confucius4_r2t2.rollback_punctuation true, false false Keep trailing text uncommitted when it already ends with punctuation instead of rolling back tokens.
confucius4_r2t2.max_tokens integer 32 Greedy decode budget per chunk and for the final flush.
confucius4_r2t2.audio_encoder_weight_type native, f32, f16 native Audio tower weight storage.
confucius4_r2t2.thinker_weight_type (alias confucius4_r2t2.weight_type) native, f32, f16, bf16, q8_0 native Thinker weight storage.
confucius4_r2t2.audio_encoder_graph_arena_mb MB 128 Audio tower graph arena.
confucius4_r2t2.thinker_prefill_graph_arena_mb MB 256 Thinker prefill graph arena.
confucius4_r2t2.thinker_decode_graph_arena_mb MB 256 Thinker decode graph arena.
confucius4_r2t2.thinker_weight_context_mb MB 64 Thinker weight context.

Request options (use with --request-option)

Option Values Default Meaning
language (or --language) language name empty Force the recognition language; Auto keeps detection.
max_tokens (or --max-tokens) integer model config Offline decode budget.

How the streaming state machine works

Each full chunk:

  1. Appends the chunk to the accumulated audio (nothing is dropped, no padding).
  2. Builds the prompt as chat template + stable prefix, where the stable prefix is the accumulated transcript minus unfixed_token_num tokens. Before unfixed_chunk_num chunks the prefix is empty. Decoding a prefix never cuts a multi-byte token: the rollback grows until the decoded prefix contains no U+FFFD replacement character.
  3. Greedy-decodes with a max_new_tokens budget on the accumulated audio.
  4. Normalizes punctuation by context, reparses the language X<asr_text>text output, and re-derives the stable prefix (again minus unfixed_token_num tokens).
  5. Reports the growth of the stable prefix as committed text.

finish_stream flushes any tail shorter than one chunk with a fixed rollback, then returns the complete transcript.

Incremental compute

Re-decoding the whole accumulated audio is the reference algorithm, but not the reference cost. Two properties keep each chunk bounded:

  • The audio tower is block-diagonal per n_window_infer attention window (4 s), with per-chunk convolutions and positions, so embeddings of completed windows never change. They are cached, and only the trailing partial window is re-encoded (together with the previous window when the tail is shorter than one conv chunk, so the reusable graph is kept).
  • The thinker prompt is template head, audio tokens, template close, stable prefix. The K/V rows for the head and the cached audio tokens are retained in the decode cache between chunks (GreedyQwenDecoderRuntime::generate_incremental), so only the tail audio, the close, and the stable text prefix are prefilled. The cache grows in 512-step buckets; crossing one falls back to a full prefill for that chunk.

Per-chunk cost is therefore dominated by the stable text prefix and the greedy tail decode, not by the utterance length.

Peak normalization and the log-mel floor are computed over all accumulated audio, so a louder segment rewrites the features of earlier windows. Cached windows are therefore reused only while their log-mel features are bit-identical to the current ones; otherwise the encoder cache and the thinker rows of those audio tokens are dropped and rebuilt, and that chunk costs a full re-encode and prefill.

Committed-text contract

Committed deltas contain transcript text only. In automatic language mode, a rolled-back prefix is not publishable until it contains <asr_text>; metadata fragments such as language neither emit a delta nor advance the published offset. Forced-language prompts already supply the metadata, so untagged decoded prefixes remain valid transcript text in that mode.

After metadata removal, growing stable prefixes emit only the suffix beyond the previously published length, counted in Unicode code points. A shrinking prefix emits nothing. The uncommitted tail is delivered in the final result (transcript.text.done on the server), rather than as a final delta.

The reference also ships a rolling-window variant for unbounded streams ("no reset": keep 16 s of audio, discard the oldest 8 s and the matching text). This port implements the standard variant used by the upstream WebSocket server, which bounds audio per utterance with VAD. Because the audio tower uses 1500 positions, a single unsegmented stream is limited to roughly 110 s of accumulated audio; segment longer streams (as the reference server does) or add the rolling-window variant.

Streaming graph reuse

Streaming no longer rebuilds the major inference graphs on every chunk. Opt-in via encode(features, reuse_graph) and R2T2ASRGenerationOptions::reuse_graphs (the streaming session enables both; the offline path keeps exact per-frame graphs):

  • The audio encoder builds one graph sized to a capacity bucket (one to two chunks exactly, then four-chunk steps) and refills the attention mask for the valid token prefix on each run, so a growing stream reuses the graph until the next bucket boundary.
  • The thinker prefill runs through the shared Qwen chunked prefill runtime (64-token blocks, one reserved graph) and the decode graph keeps a KV-cache capacity grown in 128-token buckets up to max_position_embeddings; new prompts clear KV on device.
  • Prompt token embeddings reuse a fixed-width lookup graph instead of a per-request build.

The Metal encoder avoids padding an input with fewer than 64 tokens to a capacity of 64 tokens or more, and rebuilds on shrink across that boundary. The bundled Metal backend switches attention value multiplication to a half-input SIMD-group matrix kernel at that size. A larger bucket can therefore change arithmetic precision, not just floating-point summation order.

Diagnostics on the M3 found identical convolution output, first-layer norm, and Q/K/V for the tested inputs; the first difference appeared in attention output when this boundary was crossed. Padding is not inherently inexact: some padded shapes matched bit for bit. These findings describe the tested backend and shapes, not a proof that all padding differences have one cause.

test_confucius4_r2t2_graph_reuse reads the actual retained graph capacity, including after shrink. It checks the Metal precision boundary, unpadded output equality, padded repeatability, and a backend-specific relative-RMSE drift alarm on the synthetic fixture (2e-3 on Metal under the 64-token guard, 2.5e-2 on CPU, where fp32 reduction-order noise is larger). That alarm is not a transcript-quality guarantee. Joint tests pass exact encoder output through the exact decoder and reused encoder output through the reused decoder, requiring identical token IDs for real English audio prefixes with automatic and forced-English prompts. They also check nonempty output, bucket boundaries, and growth followed by shrink. The separate decoder-only test continues to cover injected prompt lengths.

For a repeatable local benchmark, use the same checkpoint, backend, thread count, and audio on both revisions. Count positive graph.build_ms entries in --log-file output and compare confucius4_r2t2.session.stream.wall_ms, CLI deltas, final text, and process memory measurements. The bundled English sample is 14.0719 seconds (44 decode calls including the final flush at the default 320 ms chunk size). A single timing comparison is not a general performance guarantee; peak memory footprint and maximum resident set size are different measurements.

The port still recomputes the accumulated audio and prompt on each chunk; incremental encoder/prefill caching is a separate task.

Known issue: final word needs trailing silence

When the audio ends right after the last word, the model often drops that word. Push-to-talk dictation hits this when the key is released as the word ends. Both modes are affected, so it is not a streaming chunk-boundary effect. With v0.9.0, Metal, the Q8_0 GGUF and --language English, on short 16 kHz dictations:

Audio Offline Streaming (320 ms)
As recorded (0.9 s) It seems It seems
Plus 300 ms of silence It seems fixed. It seems fixed.

On the live server route, another 1.1 s clip gave Thank you very and Thank you very much. with 200 ms or more of silence. This has not been compared against the reference implementation yet.

Clients that end a stream when the user stops speaking should append about 300 ms of zero samples before closing the request body (or before EOF on --audio -).

Server usage

server.local.json declares r2t2-asr (offline) and r2t2-asr-stream (streaming, 320 ms chunks):

curl http://127.0.0.1:8488/v1/audio/transcriptions \
  -F model=r2t2-asr -F file=@speech.wav
# Decoding deltas of an uploaded file
curl -N http://127.0.0.1:8488/v1/audio/transcriptions \
  -F model=r2t2-asr-stream -F stream=true -F file=@speech.wav
# Live PCM: deltas appear while the audio is still arriving
ffmpeg -f avfoundation -i ":0" -ar 16000 -ac 1 -f s16le - \
  | curl -N -X POST -H 'Expect:' -T - \
      'http://127.0.0.1:8488/v1/audio/transcriptions/live?model=r2t2-asr-stream&sample_rate=16000&channels=1&sample_format=s16le'

Add &prompt=<url-encoded hotwords> to bias the live stream the same way --text does on the CLI.

GGUF checkpoints

GGUF is supported for f32, f16, bf16, q8_0 and the k-quants q2_k, q4_k and q6_k. f16 and q8_0 reproduce the reference transcripts exactly; storage types outside that set (legacy q4_0, q5_k, ...) are rejected at load time with an actionable error instead of silently decoding to empty text, because this graph's kernels are not validated for them.

Verified conversions are published at davidxifeng/Confucius4-R2T2-gguf (a quantized derivative work, distributed under the upstream NetEase Youdao model license — see the repository's LICENSE, LICENSE_zh, and NOTICE):

Package id File Quantization
confucius4_r2t2_q8_0 (default) r2t2-q8_0.gguf Q8_0
confucius4_r2t2_f16 r2t2-f16.gguf F16

For current file sizes, SHA-256 checksums, runtime compatibility, and published checkpoint validation, see the Hugging Face model card. Release-specific artifact details are maintained there rather than duplicated in this document.

Community Q4_K_M quantization

A community Q4_K_M quantization by NairoDorian is published at Nairod785/Confucius4-R2T2-Q4_K_M-GGUF and pinned to a fixed revision. It is a quantized derivative work under the same NetEase Youdao model license (LICENSE, LICENSE_zh, NOTICE in that repository); QUANTIZATION.md there documents the per-tensor recipe and the measured speed/quality results.

Package id File Quantization
confucius4_r2t2_q4_k_m r2t2-q4_k_m.gguf Q4_K_M (Q4_K with Q6_K MLP-down and Q2_K embeddings; BF16 audio tower)
python3 tools/model_manager_v2.py install confucius4_r2t2_q4_k_m
python3 tools/model_manager_v2.py install confucius4_r2t2_q8_0     # or confucius4_r2t2_f16

These files are self-contained: tokenizer, processor/generation config, chat template and the model spec are embedded, and the embedded spec carries the schema-v1 option contract, so option validation comes from the file itself rather than from a spec installed beside the runtime (a legacy embedded spec triggers a [warning][model_spec] fallback instead).

Converting a checkpoint yourself

Convert a real (non-symlinked) checkpoint directory:

audiocpp_gguf \
  --input /path/to/Confucius4-R2T2/model.safetensors \
  --root /path/to/Confucius4-R2T2 \
  --family confucius4_r2t2 \
  --output models/Confucius4-R2T2-GGUF/r2t2-q8_0.gguf \
  --type q8_0

The converter resolves sidecars from a real directory, so run it against a directory whose files are not symlinks (a plain cp -r of the checkpoint, or a directory created by the model manager). Direct loading of a symlinked HF cache snapshot is supported and does not need this step. The output embeds the sidecars and the model spec, so the single .gguf is portable:

audiocpp_cli --task asr --family confucius4_r2t2 \
  --model models/Confucius4-R2T2-GGUF/r2t2-q8_0.gguf --backend metal \
  --audio speech_16k.wav

Loading accepts either the .gguf file, a directory holding one, or a directory holding both formats — a GGUF checkpoint wins when both are present, matching the Qwen-family convention.

The published checkpoint ties the LM head to the token embedding and therefore contains no lm_head.weight; the family detects that and reuses the embedding, so conversions need no special flags.

Verification

The port is verified against the macOS MPS reference with golden traces:

# 1. Golden from the reference implementation (in the Confucius4-R2T2 repo)
cd /path/to/Confucius4-R2T2
PYTHONPATH=. uv run python /path/to/audio.cpp/tests/confucius4_r2t2/make_golden.py \
  --model_path /path/to/audio.cpp/models/Confucius4-R2T2 \
  --audio resources/test.wav \
  --out /path/to/audio.cpp/tests/confucius4_r2t2/golden.json

# 2. Compare the C++ runtime (offline text, per-chunk committed text, final text)
cd /path/to/audio.cpp
python3 tests/confucius4_r2t2/compare.py \
  --cli build/macos-metal-release/bin/audiocpp_cli \
  --model models/Confucius4-R2T2 \
  --audio /tmp/test.wav \
  --golden tests/confucius4_r2t2/golden.json --backend metal

compare.py runs the CLI with trace logging, parses the per-chunk trace, and diffs it against the golden. A repo-native smoke test runs both paths without the Python environment:

build/macos-metal-release/bin/test_confucius4_r2t2_transcription --backend metal

(It skips with exit code 125 when models/Confucius4-R2T2 or the audio asset is absent.)

Results on the reference machine

compare.py checks four things per golden: the offline transcript, the committed delta stream, the final streaming transcript, and every per-chunk fixed_text.

Golden Offline text Committed stream Final transcript Per-chunk fixed_text
golden.json (upstream Chinese resources/test.wav) exact exact exact 21/21 exact
golden_zh.json (same audio, --language Chinese) exact exact exact 21/21 exact
golden_sample16k.json (assets/resources/sample_16k.wav, English) exact exact exact 41/43 exact
golden_sample16k_bf16.json (same audio, bf16 reference) exact exact exact 41/43 exact

These are the original port's reference-parity results. The regression comparison now filters metadata-only prefixes from the unmodified Python goldens before checking committed deltas and per-chunk fixed_text. The C++ stream intentionally excludes the reference's language artifact; offline and final transcript expectations are unchanged.

The two English chunks that differ internally

On the 14 s English clip, chunks 37 and 42 differ in fixed_text by about one token of the trailing numeric span (22,00 vs 22,0, and a trailing you. predicted one chunk earlier). This is greedy-decoding sensitivity, not port logic:

  • both differences appear in the raw decoded text, before any R2T2 post-processing or rollback runs;
  • the reference disagrees with itself at the same order of magnitude — its fp16 and bf16 runs differ at chunks 13 and 37;
  • the difference never reaches the wire, because finish_stream publishes no delta and the final transcript is identical.

So per-chunk fixed_text equality is exact for the Chinese reference clip and stable to within one token for long English audio, with metadata artifacts excluded from committed output.