Tags: deepgram/deepgram-python-sdk
Tags
chore(main): release 7.12.0 (#798) 🤖 I have created a release *beep* *boop* --- ## [7.12.0](v7.11.0...v7.12.0) (2026-10-01) ### Features * **Flux TTS controls:** Add inline pause markers for Flux batch synthesis and IPA pronunciation overrides for Flux batch, Flux WebSocket, and Aura-2. Pronunciation is Early Access; pauses are Flux batch-only. See [Speed, Pause, Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls). ([#797](#797)) ([d2b5514](d2b5514)) * **TextBuilder migration:** Emit the escaped marker syntax Flux accepts: pauses as `\{pause:500ms\}` and pronunciations as `\{"word": "...", "pronounce": "<IPA>"\}`. The 7.11.0 forms are rejected or ignored by `/v2/speak`, so upgrade when using TextBuilder with Flux TTS. * **TextBuilder validation:** Match Flux batch limits: pauses must be 500-3000 ms in 100 ms increments, with at most eight per request. `pause()`, `from_ssml()`, and `build()` now raise `ValueError` for malformed, out-of-range, off-grid, or mixed pause-and-pronunciation controls before sending a request the API would reject. * **Listen v2 (Flux):** Add typed `Warning` responses while retaining dict-style response access during the v7 transition. ### Compatibility * **Topics and Intents:** Type the direct API `segments` response shape while retaining deprecated `results` facades, legacy import paths, and legacy construction throughout v7. --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: Corey Weathers <corey.weathers@deepgram.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
chore(main): release 7.11.0 (#795) 🤖 I have created a release *beep* *boop* --- ## [7.11.0](v7.10.0...v7.11.0) (2026-09-23) ### Features * **Listen v2 (Flux):** change numerals during an active connection with `connection.send_configure(ListenV2Configure(numerals=True))`. Numerals convert written numbers to numerical format for turns transcribed after the update. ([#794](#794)) ([c7c1b3f](c7c1b3f)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>
chore(main): release 7.10.0 (#791) 🤖 I have created a release *beep* *boop* --- ## [7.10.0](v7.9.0...v7.10.0) (2026-09-17) ### Features * **Voice Agent:** Add typed `FunctionCallCancelled` WebSocket events. Events identify function calls that are no longer valid; do not send a `FunctionCallResponse` for the listed call IDs. ([#790](#790)) ([7f5b642](7f5b642)) * **Voice Agent:** Add the optional `defer_until_eot` function setting. When enabled, it holds a function call until the user's turn is confirmed and discards the deferred call if the turn resumes. Defaults to `false`. ([#790](#790)) ([7f5b642](7f5b642)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>
chore(main): release 7.9.0 (#787) 🤖 I have created a release *beep* *boop* --- ## [7.9.0](v7.8.1...v7.9.0) (2026-09-14) ### Features * **Voice Agent:** Adds `send_force_end_turn()` and typed ForceEndTurn messages for V2 Flux listen providers. ([#783](#783)) ([ff49864](ff49864)) * **Voice Agent:** Adds integer `expressivity` from `-2` through `2` for Deepgram Flux TTS providers. ([#783](#783)) ([ff49864](ff49864)) * **Flux TTS:** `speed` accepts `0.5` through `1.5` in `0.05` increments, replacing the previously documented `0.85` through `1.15` range. ([#783](#783)) ([ff49864](ff49864)) ### Model Catalog * **Speak V1:** Removes `aura-2-perseo-it` from model literals. The API never served the model and returns 400 `No such model/version combination found` for it. ([#783](#783)) ([ff49864](ff49864)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>
chore(main): release 7.8.1 (#777) 🤖 I have created a release *beep* *boop* --- ## [7.8.1](v7.8.0...v7.8.1) (2026-09-03) ### Bug Fixes * **TextBuilder:** `ssml_to_deepgram()` now preserves a `<phoneme>` pronunciation when its valid `ph` and `alphabet` attributes appear in either order. ([#741](#741)) ([7fd4b63](7fd4b63)) * **Credentials:** Explicitly passing `api_key=None` continues to disable ambient `DEEPGRAM_API_KEY` lookup, which is important for multi-tenant and test environments. ([#778](#778)) ([e675990](e675990)) * **Credentials:** `DeepgramClient()` and `AsyncDeepgramClient()` now resolve `DEEPGRAM_API_KEY` when constructed, so `load_dotenv()` can run after importing the SDK. Closes [#734](#734). ([#767](#767)) ([ec362ec](ec362ec)) * **Custom transports:** Speak V2 WebSocket connections now honor `transport_factory`, matching the routing behavior of other WebSocket APIs for proxies, test doubles, and custom-hosted transports. ([#766](#766)) ([0980663](0980663)) ### Documentation * **Transcription:** Clarified that Nova-3 assumes English when `language` is omitted; non-English and multilingual audio require an explicit language such as `fr` or `multi`. ([#771](#771)) ([4574337](4574337)) * **Examples:** Added Listen V1 live microphone transcription with optional `sounddevice`, device selection, bounded audio buffering, transcript output, and clean Ctrl-C shutdown. ([#780](#780)) ([08f0471](08f0471)) * **Examples:** Added a resilient Listen V1 live transcription pattern with exponential backoff, reconnect-aware audio buffering, timestamp continuity, and clean shutdown. ([#776](#776)) ([96b2d11](96b2d11)) * **Examples:** Added an application-owned Voice Agent session recorder that serializes received transcripts, function calls, and latency reports as JSON while leaving consent, redaction, retention, and storage policy to the application. Closes [#775](#775). ([#781](#781)) ([30ad152](30ad152)) * **Text-to-Speech:** Corrected streaming synthesis snippets to iterate the response byte chunks instead of accessing a nonexistent `.stream` attribute. ([#749](#749)) ([178724e](178724e)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>
chore(main): release 7.8.0 (#774) 🤖 I have created a release *beep* *boop* --- ## [7.8.0](v7.7.1...v7.8.0) (2026-08-28) ### Features * **regen:** listen v2 force-end-turn and listen v1 diarize metadata ([#768](#768)) ([bacd1b5](bacd1b5)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
chore(main): release 7.7.1 (#770) 🤖 I have created a release *beep* *boop* --- ## [7.7.1](v7.7.0...v7.7.1) (2026-08-21) ### Bug Fixes * **listen/v2:** restore subscript access on typed responses ([c4d2580](c4d2580)) * narrow listen v2 compatibility to subscript access ([aac5e0a](aac5e0a)) * preserve dict access on typed response models ([b9f544a](b9f544a)) * restore listen v2 mapping behavior ([8a18e3b](8a18e3b)) * scope dict compatibility to listen v2 responses ([0c67f00](0c67f00)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
chore(main): release 7.7.0 (#764) ## 7.7.0 Flux TTS streaming controls and Listen v2 redaction. ### Features **Speak v2 (Flux TTS streaming)** - **Barge-in** — `send_interrupt()` (optionally with a `playback_offset`, `{type: "time_ms", value: N}`), answered by a `SpeechInterrupted` server message whose metadata carries the new `controls_applied.breaks_applied` counter. - **Mid-stream `send_configure()`** — change `speed` mid-session, acknowledged by `ConfigureSuccess` / `ConfigureFailure`. - **`speed`** and **`expressivity`** connect query parameters. - Inline pause and pronunciation controls are not applied at launch — they are stripped before synthesis and support is coming soon. **Listen v2** - **`redact`** connect parameter (`ListenV2Redact`: `numbers`, `aggressive_numbers`). - `send_configure()` is now properly typed (`ListenV2Configure` + `ListenV2ConfigureSuccess` in the response union), replacing the previous `typing.Any` shim. **Other** - `GoogleThinkProviderVersion` adds `ai-studio-v1beta` and `gemini-enterprise-agent-v1`. - `AgentV1UpdateListenListenProvider` discriminated union (`_V1` / `_V2`, discriminant `version`). - `client_wrapper` now derives its version from `importlib.metadata` rather than a hardcoded string. ### Compatibility No breaking changes against v7.6.0: 0 removed public exports, 0 deleted modules, baseline socket-client signatures intact, and enum changes are widenings only. - The `deepgram` speak provider `version` widens from `Literal["v1"]` to `str`. - `AgentV1UpdateListenListen.provider` moves from a bare `DeepgramListenProviderV2` to a required discriminated union; a compatibility validator coerces a legacy provider instance or bare dict into the new shape (both serialize to `version: "v2"`), so existing callers are unaffected. --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>
chore(main): release 7.6.0 (#747) 🤖 I have created a release *beep* *boop* --- ## [7.6.0](v7.5.0...v7.6.0) (2026-07-22) ### Features * **regen:** flux stt numerals and aura-2 multilingual tts voices ([#746](#746)) ([17a1deb](17a1deb)) #### What's in this release SDK regeneration for **2026-07-20** (Fern generator `5.14.18`, spec `ff8fd2b7`). **Flux STT `numerals`** - New `ListenV2Numerals` type (`Literal["true", "false"]`) in `types/listen_v2numerals.py`, exported from the package root and `types`. - New `numerals` query param on `listen.v2.connect()` and the raw client (connection-time only). See [docs #1020]. - Demonstrated in `examples/14-transcription-live-websocket-v2.py`. - New wire coverage in `tests/custom/test_listen_v2_connect_wire.py` (sync + async) pins that `/v2/listen` serializes `numerals=true` and omits it when absent — the `/v2/listen` handshake previously had no wire test. **New Aura-2 multilingual TTS voices** - ~40 voices added to `SpeakV1Model` and `AudioGenerateRequestModel` (es / de / nl / fr / it / ja). - Additive only — not breaking. #### Breaking change from spec — mitigated (stays non-breaking) - The spec removed `stt_latency` from `AgentV1LatencyReport` (docs #1006). Because `AgentV1LatencyReport` is a server-emitted (read-only) message, `stt_latency` is re-added by hand as a read-side backward-compat field (frozen in `.fernignore`): `report.stt_latency` keeps resolving to `None` instead of raising `AttributeError`, with **no request/wire impact**. Covered by `tests/custom/test_latency_report_stt_compat.py`. This keeps the release a minor bump rather than a major. #### Verification - `mypy src/deepgram` — clean (857 files) - `pytest` — 331 passed, 1 skipped - `ruff check src/deepgram` — clean --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
chore(main): release 7.5.0 (#743) Release PR for **7.5.0** — [compare v7.4.0...v7.5.0](v7.4.0...v7.5.0) ## What's in 7.5.0 - **[#742](#742) — Flux TTS streaming (Speak V2) + agent listen reconfigure** - **Speak V2 streaming (WebSocket)** — `client.speak.v2.connect(...)`: new `speak/v2/` package (client, socket client, request/response types), top-level `SpeakV2Encoding` / `SpeakV2MipOptOut` / `SpeakV2Model` / `SpeakV2SampleRate` / `SpeakV2Tag`, and optional `send_flush()` / `send_close()` control messages. - **Agent mid-session listen reconfigure** — `send_update_listen(...)` with `AgentV1UpdateListen` / `AgentV1UpdateListenListen`, plus the `AgentV1ListenUpdated` server acknowledgement. - **Flux end-of-turn tuning** — new listen-provider fields `eot_threshold`, `eager_eot_threshold`, `eot_timeout_ms`. - **[#744](#744) — Flux TTS batch (REST) + agent latency report** - **Speak V2 batch (REST)** — `client.speak.v2.audio.generate(...)` (+ raw client): the REST companion to streaming. `model` required (flux-only); optional `encoding` / `container` / `sample_rate` / `bit_rate` / `callback` / `callback_method` / `tag` / `mip_opt_out` / `priority`. No `speed` (GA-only, excluded for EA). - `sample_rate` / `bit_rate` are typed as `int` — a float like `24000.0` is rejected on the wire. - Callback mode returns the `SpeakV2AcceptedResponse` JSON ack through the audio byte iterator as raw bytes; join the chunks and parse `request_id` yourself (documented in the docstring; typed content-type dispatch tracked as a follow-up). - **`AgentV1LatencyReport`** — new latency-report type + request, wired into the agent socket-client response union. - **Agent inject-message `interrupt`** — new value on `AgentV1InjectAgentMessageBehavior`. --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>
PreviousNext