Skip to content

Latest commit

 

History

History
133 lines (115 loc) · 37 KB

File metadata and controls

133 lines (115 loc) · 37 KB

Router.Audio.Transcriptions

Overview

Available Operations

  • create - Create transcription

create

Transcribe audio input into text using the configured transcription model and return the result.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.router.audio.transcriptions.create(model="Malibu", enable_logging=True, diarize=False, tag_audio_events=True, timestamps_granularity="word", temperature=0.5, timestamp_granularities=[
        "word",
        "segment",
    ], retry={
        "on_codes": [
            429.0,
            500.0,
            502.0,
            503.0,
            504.0,
        ],
    }, load_balancer={
        "type": "weight_based",
        "models": [
            {
                "model": "openai/gpt-4o",
                "weight": 0.7,
            },
        ],
    }, timeout={
        "call_timeout": 30000.0,
    }, orq={
        "fallbacks": [
            {
                "model": "openai/gpt-4o-mini",
            },
        ],
        "retry": {
            "on_codes": [
                429.0,
                500.0,
                502.0,
                503.0,
                504.0,
            ],
        },
        "identity": {
            "id": "contact_01ARZ3NDEKTSV4RRFFQ69G5FAV",
            "display_name": "Jane Doe",
            "email": "jane.doe@example.com",
            "metadata": [
                {
                    "department": "Engineering",
                    "role": "Senior Developer",
                },
            ],
            "logo_url": "https://example.com/avatars/jane-doe.jpg",
            "tags": [
                "hr",
                "engineering",
            ],
        },
        "load_balancer": {
            "type": "weight_based",
            "models": [
                {
                    "model": "openai/gpt-4o",
                    "weight": 0.7,
                },
                {
                    "model": "anthropic/claude-3-5-sonnet",
                    "weight": 0.3,
                },
            ],
        },
        "timeout": {
            "call_timeout": 30000.0,
        },
    })

    # Handle response
    print(res)

Parameters

Parameter Type Required Description Example
model str ✔️ ID of the model to use openai/gpt-4o-transcribe
prompt Optional[str] ➖ An optional text to guide the model's style or continue a previous audio segment. The prompt should match the audio language.
enable_logging Optional[bool] ➖ When enable_logging is set to false, zero retention mode is used. This disables history features like request stitching and is only available to enterprise customers.
diarize Optional[bool] ➖ Whether to annotate which speaker is currently talking in the uploaded file.
response_format Optional[models.CreateTranscriptionResponseFormat] ➖ The format of the transcript output, in one of these options: json, text, srt, verbose_json, or vtt.
tag_audio_events Optional[bool] ➖ Whether to tag audio events like (laughter), (footsteps), etc. in the transcription.
num_speakers Optional[float] ➖ The maximum amount of speakers talking in the uploaded file. Helps with predicting who speaks when, the maximum is 32.
timestamps_granularity Optional[models.TimestampsGranularity] ➖ The granularity of the timestamps in the transcription. Word provides word-level timestamps and character provides character-level timestamps per word.
temperature Optional[float] ➖ The sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. If set to 0, the model will use log probability to automatically increase the temperature until certain thresholds are hit. 0.5
language Optional[str] ➖ The language of the input audio. Supplying the input language in ISO-639-1 format will improve accuracy and latency.
timestamp_granularities List[models.TimestampGranularities] ➖ The timestamp granularities to populate for this transcription. response_format must be set to verbose_json to use timestamp granularities. Either or both of these options are supported: "word" or "segment". Note: There is no additional latency for segment timestamps, but generating word timestamps incurs additional latency. [
"word",
"segment"
]
name Optional[str] ➖ The name to display on the trace. If not specified, the default system name will be used.
fallbacks List[models.CreateTranscriptionFallbacks] ➖ Array of fallback models to use if primary model fails
retry Optional[models.CreateTranscriptionRetry] ➖ Retry configuration for the request
load_balancer Optional[models.CreateTranscriptionLoadBalancer] ➖ Load balancer configuration for the request.
timeout Optional[models.CreateTranscriptionTimeout] ➖ Timeout configuration to apply to the request. If the request exceeds the timeout, it will be retried or fallback to the next model if configured.
orq Optional[models.CreateTranscriptionOrq] ➖ N/A
file Optional[models.CreateTranscriptionFile] ➖ The audio file object (not file name) to transcribe, in one of these formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, or webm.
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.CreateTranscriptionResponseBody

Errors

Error Type Status Code Content Type
models.CreateTranscriptionRouterAudioTranscriptionsResponseBody 422 application/json
models.APIDefaultError 4XX, 5XX */*