Skip to content

Tags: deepgram/deepgram-python-sdk

Tags

v7.12.0

Toggle v7.12.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.12.0 (#798)

🤖 I have created a release *beep* *boop*

---

##
[7.12.0](v7.11.0...v7.12.0)
(2026-10-01)

### Features

* **Flux TTS controls:** Add inline pause markers for Flux batch
synthesis and IPA pronunciation overrides for Flux batch, Flux
WebSocket, and Aura-2. Pronunciation is Early Access; pauses are Flux
batch-only. See [Speed, Pause,
Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls).
([#797](#797))
([d2b5514](d2b5514))
* **TextBuilder migration:** Emit the escaped marker syntax Flux
accepts: pauses as `\{pause:500ms\}` and pronunciations as `\{"word":
"...", "pronounce": "<IPA>"\}`. The 7.11.0 forms are rejected or ignored
by `/v2/speak`, so upgrade when using TextBuilder with Flux TTS.
* **TextBuilder validation:** Match Flux batch limits: pauses must be
500-3000 ms in 100 ms increments, with at most eight per request.
`pause()`, `from_ssml()`, and `build()` now raise `ValueError` for
malformed, out-of-range, off-grid, or mixed pause-and-pronunciation
controls before sending a request the API would reject.
* **Listen v2 (Flux):** Add typed `Warning` responses while retaining
dict-style response access during the v7 transition.

### Compatibility

* **Topics and Intents:** Type the direct API `segments` response shape
while retaining deprecated `results` facades, legacy import paths, and
legacy construction throughout v7.

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: Corey Weathers <corey.weathers@deepgram.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

v7.11.0

Toggle v7.11.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.11.0 (#795)

🤖 I have created a release *beep* *boop*
---


##
[7.11.0](v7.10.0...v7.11.0)
(2026-09-23)


### Features

* **Listen v2 (Flux):** change numerals during an active connection with
`connection.send_configure(ListenV2Configure(numerals=True))`. Numerals
convert written numbers to numerical format for turns transcribed after
the update.
([#794](#794))
([c7c1b3f](c7c1b3f))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>

v7.10.0

Toggle v7.10.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.10.0 (#791)

🤖 I have created a release *beep* *boop*
---

##
[7.10.0](v7.9.0...v7.10.0)
(2026-09-17)

### Features

* **Voice Agent:** Add typed `FunctionCallCancelled` WebSocket events.
Events identify function calls that are no longer valid; do not send a
`FunctionCallResponse` for the listed call IDs.
([#790](#790))
([7f5b642](7f5b642))
* **Voice Agent:** Add the optional `defer_until_eot` function setting.
When enabled, it holds a function call until the user's turn is
confirmed and discards the deferred call if the turn resumes. Defaults
to `false`.
([#790](#790))
([7f5b642](7f5b642))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>

v7.9.0

Toggle v7.9.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.9.0 (#787)

🤖 I have created a release *beep* *boop*
---

##
[7.9.0](v7.8.1...v7.9.0)
(2026-09-14)

### Features

* **Voice Agent:** Adds `send_force_end_turn()` and typed ForceEndTurn
messages for V2 Flux listen providers.
([#783](#783))
([ff49864](ff49864))
* **Voice Agent:** Adds integer `expressivity` from `-2` through `2` for
Deepgram Flux TTS providers.
([#783](#783))
([ff49864](ff49864))
* **Flux TTS:** `speed` accepts `0.5` through `1.5` in `0.05`
increments, replacing the previously documented `0.85` through `1.15`
range.
([#783](#783))
([ff49864](ff49864))

### Model Catalog

* **Speak V1:** Removes `aura-2-perseo-it` from model literals. The API
never served the model and returns 400 `No such model/version
combination found` for it.
([#783](#783))
([ff49864](ff49864))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>

v7.8.1

Toggle v7.8.1's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.8.1 (#777)

🤖 I have created a release *beep* *boop*
---


##
[7.8.1](v7.8.0...v7.8.1)
(2026-09-03)


### Bug Fixes

* **TextBuilder:** `ssml_to_deepgram()` now preserves a `<phoneme>`
pronunciation when its valid `ph` and `alphabet` attributes appear in
either order.
([#741](#741))
([7fd4b63](7fd4b63))
* **Credentials:** Explicitly passing `api_key=None` continues to
disable ambient `DEEPGRAM_API_KEY` lookup, which is important for
multi-tenant and test environments.
([#778](#778))
([e675990](e675990))
* **Credentials:** `DeepgramClient()` and `AsyncDeepgramClient()` now
resolve `DEEPGRAM_API_KEY` when constructed, so `load_dotenv()` can run
after importing the SDK. Closes
[#734](#734).
([#767](#767))
([ec362ec](ec362ec))
* **Custom transports:** Speak V2 WebSocket connections now honor
`transport_factory`, matching the routing behavior of other WebSocket
APIs for proxies, test doubles, and custom-hosted transports.
([#766](#766))
([0980663](0980663))


### Documentation

* **Transcription:** Clarified that Nova-3 assumes English when
`language` is omitted; non-English and multilingual audio require an
explicit language such as `fr` or `multi`.
([#771](#771))
([4574337](4574337))
* **Examples:** Added Listen V1 live microphone transcription with
optional `sounddevice`, device selection, bounded audio buffering,
transcript output, and clean Ctrl-C shutdown.
([#780](#780))
([08f0471](08f0471))
* **Examples:** Added a resilient Listen V1 live transcription pattern
with exponential backoff, reconnect-aware audio buffering, timestamp
continuity, and clean shutdown.
([#776](#776))
([96b2d11](96b2d11))
* **Examples:** Added an application-owned Voice Agent session recorder
that serializes received transcripts, function calls, and latency
reports as JSON while leaving consent, redaction, retention, and storage
policy to the application. Closes
[#775](#775).
([#781](#781))
([30ad152](30ad152))
* **Text-to-Speech:** Corrected streaming synthesis snippets to iterate
the response byte chunks instead of accessing a nonexistent `.stream`
attribute.
([#749](#749))
([178724e](178724e))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>

v7.8.0

Toggle v7.8.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.8.0 (#774)

🤖 I have created a release *beep* *boop*
---


##
[7.8.0](v7.7.1...v7.8.0)
(2026-08-28)


### Features

* **regen:** listen v2 force-end-turn and listen v1 diarize metadata
([#768](#768))
([bacd1b5](bacd1b5))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

v7.7.1

Toggle v7.7.1's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.7.1 (#770)

🤖 I have created a release *beep* *boop*
---


##
[7.7.1](v7.7.0...v7.7.1)
(2026-08-21)


### Bug Fixes

* **listen/v2:** restore subscript access on typed responses
([c4d2580](c4d2580))
* narrow listen v2 compatibility to subscript access
([aac5e0a](aac5e0a))
* preserve dict access on typed response models
([b9f544a](b9f544a))
* restore listen v2 mapping behavior
([8a18e3b](8a18e3b))
* scope dict compatibility to listen v2 responses
([0c67f00](0c67f00))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

v7.7.0

Toggle v7.7.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.7.0 (#764)

## 7.7.0

Flux TTS streaming controls and Listen v2 redaction.

### Features

**Speak v2 (Flux TTS streaming)**
- **Barge-in** — `send_interrupt()` (optionally with a
`playback_offset`, `{type: "time_ms", value: N}`), answered by a
`SpeechInterrupted` server message whose metadata carries the new
`controls_applied.breaks_applied` counter.
- **Mid-stream `send_configure()`** — change `speed` mid-session,
acknowledged by `ConfigureSuccess` / `ConfigureFailure`.
- **`speed`** and **`expressivity`** connect query parameters.
- Inline pause and pronunciation controls are not applied at launch —
they are stripped before synthesis and support is coming soon.

**Listen v2**
- **`redact`** connect parameter (`ListenV2Redact`: `numbers`,
`aggressive_numbers`).
- `send_configure()` is now properly typed (`ListenV2Configure` +
`ListenV2ConfigureSuccess` in the response union), replacing the
previous `typing.Any` shim.

**Other**
- `GoogleThinkProviderVersion` adds `ai-studio-v1beta` and
`gemini-enterprise-agent-v1`.
- `AgentV1UpdateListenListenProvider` discriminated union (`_V1` /
`_V2`, discriminant `version`).
- `client_wrapper` now derives its version from `importlib.metadata`
rather than a hardcoded string.

### Compatibility

No breaking changes against v7.6.0: 0 removed public exports, 0 deleted
modules, baseline socket-client signatures intact, and enum changes are
widenings only.
- The `deepgram` speak provider `version` widens from `Literal["v1"]` to
`str`.
- `AgentV1UpdateListenListen.provider` moves from a bare
`DeepgramListenProviderV2` to a required discriminated union; a
compatibility validator coerces a legacy provider instance or bare dict
into the new shape (both serialize to `version: "v2"`), so existing
callers are unaffected.

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>

v7.6.0

Toggle v7.6.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.6.0 (#747)

🤖 I have created a release *beep* *boop*
---


##
[7.6.0](v7.5.0...v7.6.0)
(2026-07-22)


### Features

* **regen:** flux stt numerals and aura-2 multilingual tts voices
([#746](#746))
([17a1deb](17a1deb))

#### What's in this release

SDK regeneration for **2026-07-20** (Fern generator `5.14.18`, spec
`ff8fd2b7`).

**Flux STT `numerals`**
- New `ListenV2Numerals` type (`Literal["true", "false"]`) in
`types/listen_v2numerals.py`, exported from the package root and
`types`.
- New `numerals` query param on `listen.v2.connect()` and the raw client
(connection-time only). See [docs #1020].
- Demonstrated in `examples/14-transcription-live-websocket-v2.py`.
- New wire coverage in `tests/custom/test_listen_v2_connect_wire.py`
(sync + async) pins that `/v2/listen` serializes `numerals=true` and
omits it when absent — the `/v2/listen` handshake previously had no wire
test.

**New Aura-2 multilingual TTS voices**
- ~40 voices added to `SpeakV1Model` and `AudioGenerateRequestModel` (es
/ de / nl / fr / it / ja).
- Additive only — not breaking.

#### Breaking change from spec — mitigated (stays non-breaking)

- The spec removed `stt_latency` from `AgentV1LatencyReport` (docs
#1006). Because `AgentV1LatencyReport` is a server-emitted (read-only)
message, `stt_latency` is re-added by hand as a read-side
backward-compat field (frozen in `.fernignore`): `report.stt_latency`
keeps resolving to `None` instead of raising `AttributeError`, with **no
request/wire impact**. Covered by
`tests/custom/test_latency_report_stt_compat.py`. This keeps the release
a minor bump rather than a major.

#### Verification

- `mypy src/deepgram` — clean (857 files)
- `pytest` — 331 passed, 1 skipped
- `ruff check src/deepgram` — clean

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

v7.5.0

Toggle v7.5.0's commit message

Verified

This commit was created on GitHub.com and signed with GitHub’s verified signature.
chore(main): release 7.5.0 (#743)

Release PR for **7.5.0** — [compare
v7.4.0...v7.5.0](v7.4.0...v7.5.0)

## What's in 7.5.0

- **[#742](#742) —
Flux TTS streaming (Speak V2) + agent listen reconfigure**
- **Speak V2 streaming (WebSocket)** — `client.speak.v2.connect(...)`:
new `speak/v2/` package (client, socket client, request/response types),
top-level `SpeakV2Encoding` / `SpeakV2MipOptOut` / `SpeakV2Model` /
`SpeakV2SampleRate` / `SpeakV2Tag`, and optional `send_flush()` /
`send_close()` control messages.
- **Agent mid-session listen reconfigure** — `send_update_listen(...)`
with `AgentV1UpdateListen` / `AgentV1UpdateListenListen`, plus the
`AgentV1ListenUpdated` server acknowledgement.
- **Flux end-of-turn tuning** — new listen-provider fields
`eot_threshold`, `eager_eot_threshold`, `eot_timeout_ms`.

- **[#744](#744) —
Flux TTS batch (REST) + agent latency report**
- **Speak V2 batch (REST)** — `client.speak.v2.audio.generate(...)` (+
raw client): the REST companion to streaming. `model` required
(flux-only); optional `encoding` / `container` / `sample_rate` /
`bit_rate` / `callback` / `callback_method` / `tag` / `mip_opt_out` /
`priority`. No `speed` (GA-only, excluded for EA).
- `sample_rate` / `bit_rate` are typed as `int` — a float like `24000.0`
is rejected on the wire.
- Callback mode returns the `SpeakV2AcceptedResponse` JSON ack through
the audio byte iterator as raw bytes; join the chunks and parse
`request_id` yourself (documented in the docstring; typed content-type
dispatch tracked as a follow-up).
- **`AgentV1LatencyReport`** — new latency-report type + request, wired
into the agent socket-client response union.
- **Agent inject-message `interrupt`** — new value on
`AgentV1InjectAgentMessageBehavior`.

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please).

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Greg Holmes <greg.holmes@deepgram.com>