Skip to content

About

Generic realtime voice transcription bridge from telephony media streams to realtime model sessions and signed downstream webhook events.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Realtime Voice Transcription Bridge Reference

Generic reference bridge for taking a live telephony media stream, forwarding audio into an OpenAI-compatible realtime model session, and emitting signed downstream webhook events.

It is intentionally not tied to one CRM, one domain, or one extraction schema. Twilio is included as the first telephony adapter, but the internal bridge contracts are provider-neutral so teams can add SIP, browser WebRTC, contact-center, or recording-file adapters later.

Reference call flow

Inbound call
  -> telephony webhook returns stream instructions
  -> media WebSocket connects
  -> realtime model session opens
  -> audio chunks stream to model
  -> transcript events arrive
  -> structured field deltas arrive
  -> bridge accumulates safe state
  -> signed downstream webhook events are delivered
  -> call/session closes cleanly

Why this matters in real realtime systems

  • Telephony streams and model sessions are two separate WebSocket lifecycles; either side can close first.
  • Live extraction should emit deltas, not wait for the whole call, so downstream systems can update while the conversation is happening.
  • Downstream event receivers need signatures and idempotent event IDs because realtime bridges are public-facing integration points.
  • The extraction schema should be configuration, not proprietary code, so the bridge can be reused across intake, support, scheduling, triage, and QA workflows.

What this repo demonstrates

  • Node.js HTTP server with a Twilio Media Streams adapter.
  • Provider-neutral bridge session contract.
  • OpenAI-compatible realtime WebSocket client using configurable URL, model, audio format, transcription model, tool name, and JSON schema.
  • Generic extraction schema loading from JSON.
  • Signed downstream webhook delivery with HMAC SHA-256.
  • Event envelopes for call.started, transcript.final, field.delta, call.ended, and bridge.error.
  • Tests for TwiML generation, media event parsing, model event mapping, state accumulation, webhook signatures, and bridge orchestration.

Local setup

Install dependencies:

npm install

Run checks:

npm test
npm run build
npm run lint

Run locally:

cp .env.example .env
npm run dev

Expose the local server to your telephony provider with a tunnel such as ngrok or Cloudflare Tunnel, then set:

PUBLIC_BASE_URL=https://voice-bridge.example.com

Configuration

PORT=8080
PUBLIC_BASE_URL=https://voice-bridge.example.com
STREAM_PATH=/streams/twilio
REALTIME_WS_URL=wss://api.openai.com/v1/realtime?model=<model>
REALTIME_API_KEY=<secret>
REALTIME_BETA_HEADER=realtime=v1
REALTIME_INPUT_AUDIO_FORMAT=g711_ulaw
REALTIME_TRANSCRIPTION_MODEL=whisper-1
EXTRACTION_SCHEMA_PATH=examples/contact-intake.schema.json
EXTRACTION_TOOL_NAME=emit_structured_update
DOWNSTREAM_WEBHOOK_URL=https://hooks.example.com/realtime-events
DOWNSTREAM_SIGNING_SECRET=<secret>

Twilio adapter

Configure your number voice webhook:

POST https://voice-bridge.example.com/voice/twilio

The endpoint returns TwiML that starts a media stream to:

wss://voice-bridge.example.com/streams/twilio?callId=<call-id>&from=<caller>

Downstream events

Every event is signed when DOWNSTREAM_SIGNING_SECRET is configured:

x-reference-timestamp: <unix-seconds>
x-reference-signature: sha256=<hex-hmac>

The HMAC message is:

<timestamp>.<json-body>

Example event:

{
  "id": "evt_...",
  "type": "field.delta",
  "callId": "call_...",
  "sequence": 4,
  "occurredAt": "2026-08-21T12:00:00.000Z",
  "payload": {
    "delta": {
      "preferred_contact_time": "morning"
    },
    "state": {
      "preferred_contact_time": "morning"
    }
  }
}

Bring your own schema

The bridge loads a JSON schema from EXTRACTION_SCHEMA_PATH and gives it to the realtime model as the parameter schema for the configured extraction tool. The included examples/contact-intake.schema.json is deliberately generic.

To adapt this repo:

  1. Replace the example schema with your own allowed fields.
  2. Update the extraction instructions in src/realtime/instructions.ts if needed.
  3. Point DOWNSTREAM_WEBHOOK_URL at your workflow, CRM, queue, or event bus.
  4. Verify webhook signatures in the receiver.

About

Generic realtime voice transcription bridge from telephony media streams to realtime model sessions and signed downstream webhook events.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages