Skip to content

Latest commit

 

History

History
91 lines (73 loc) · 82 KB

File metadata and controls

91 lines (73 loc) · 82 KB

Router.Completions

Overview

Available Operations

create

For sending requests to legacy completion models

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.router.completions.create(model="XC90", prompt="<value>", echo=False, frequency_penalty=0.0, max_tokens=16, presence_penalty=0.0, temperature=1.0, top_p=1.0, n=1, retry={
        "on_codes": [
            429.0,
            500.0,
            502.0,
            503.0,
            504.0,
        ],
    }, cache={
        "ttl": 3600.0,
        "type": "exact_match",
    }, load_balancer={
        "type": "weight_based",
        "models": [
            {
                "model": "openai/gpt-4o",
                "weight": 0.7,
            },
        ],
    }, timeout={
        "call_timeout": 30000.0,
    }, stream=False)

    with res as event_stream:
        for event in event_stream:
            # handle event
            print(event, flush=True)

Parameters

Parameter Type Required Description Example
model str ✔️ ID of the model to use openai/gpt-3.5-turbo-instruct
prompt str ✔️ The prompt(s) to generate completions for, encoded as a string, array of strings, array of tokens, or array of token arrays.
echo OptionalNullable[bool] ➖ Echo back the prompt in addition to the completion
frequency_penalty OptionalNullable[float] ➖ Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.
max_tokens OptionalNullable[int] ➖ The maximum number of tokens that can be generated in the completion.
presence_penalty OptionalNullable[float] ➖ Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.
seed OptionalNullable[int] ➖ If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result.
stop OptionalNullable[models.CreateCompletionStop] ➖ Up to 4 sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence.
temperature OptionalNullable[float] ➖ What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
top_p OptionalNullable[float] ➖ An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered.
n OptionalNullable[int] ➖ How many completions to generate for each prompt. Note: Because this parameter generates many completions, it can quickly consume your token quota.
user Optional[str] ➖ A unique identifier representing your end-user, which can help OpenAI to monitor and detect abuse.
name Optional[str] ➖ The name to display on the trace. If not specified, the default system name will be used.
fallbacks List[models.CreateCompletionFallbacks] ➖ Array of fallback models to use if primary model fails
retry Optional[models.CreateCompletionRetry] ➖ Retry configuration for the request
cache Optional[models.CreateCompletionCache] ➖ Cache configuration for the request.
load_balancer Optional[models.CreateCompletionLoadBalancer] ➖ Load balancer configuration for the request.
timeout Optional[models.CreateCompletionTimeout] ➖ Timeout configuration to apply to the request. If the request exceeds the timeout, it will be retried or fallback to the next model if configured.
thinking OptionalNullable[models.CreateCompletionThinking] ➖ Configuration for the thinking mode capability. Set type to adaptive for models that support adaptive thinking (e.g. Claude Opus 4.6, Sonnet 4.6), or enabled with budget_tokens for manual control.
plugins List[models.PIIRedactionPlugin] ➖ Request-scoped transforms applied to the text exchanged with the model. Currently supports pii_redaction, which replaces PII with placeholders before the provider sees it and restores the original values in the response.
orq Optional[models.CreateCompletionOrq] ➖ : warning: ** DEPRECATED **: This will be removed in a future release, please migrate away from it as soon as possible.

Leverage Orq's intelligent routing capabilities to enhance your AI application with enterprise-grade reliability and observability. Orq provides automatic request management including retries on failures, model fallbacks for high availability, identity-level analytics tracking, conversation threading, and dynamic prompt templating with variable substitution.
{
"retry": {
"count": 3,
"on_codes": [
429,
500,
502
]
},
"fallbacks": [
{
"model": "openai/gpt-5"
},
{
"model": "anthropic/claude-4-opus"
}
],
"identity": {
"id": "identity_01ARZ3NDEKTSV4RRFFQ69G5FAV",
"display_name": "Jane Doe",
"email": "jane.doe@example.com"
},
"thread": {
"id": "thread_01ARZ3NDEKTSV4RRFFQ69G5FAV",
"tags": [
"customer-support"
]
},
"inputs": {
"customer_name": "John Smith",
"issue_type": "billing"
},
"cache": {
"ttl": 3600,
"type": "exact_match"
},
"knowledge_bases": [
{
"knowledge_id": "knowledge_01ARZ3NDEKTSV4RRFFQ69G5FAV",
"top_k": 5
}
],
"timeout": {
"call_timeout": 30000
},
"security": {
"mask": [
"input",
"system"
]
}
}
stream Optional[bool] ➖ N/A
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.CreateCompletionResponse

Errors

Error Type Status Code Content Type
models.APIDefaultError 4XX, 5XX */*