- create - Create completion
For sending requests to legacy completion models
from orq_ai_sdk import Orq
import os
with Orq(
api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:
res = orq.router.completions.create(model="XC90", prompt="<value>", echo=False, frequency_penalty=0.0, max_tokens=16, presence_penalty=0.0, temperature=1.0, top_p=1.0, n=1, retry={
"on_codes": [
429.0,
500.0,
502.0,
503.0,
504.0,
],
}, cache={
"ttl": 3600.0,
"type": "exact_match",
}, load_balancer={
"type": "weight_based",
"models": [
{
"model": "openai/gpt-4o",
"weight": 0.7,
},
],
}, timeout={
"call_timeout": 30000.0,
}, stream=False)
with res as event_stream:
for event in event_stream:
# handle event
print(event, flush=True)| Parameter | Type | Required | Description | Example |
|---|---|---|---|---|
model |
str | ✔️ | ID of the model to use | openai/gpt-3.5-turbo-instruct |
prompt |
str | ✔️ | The prompt(s) to generate completions for, encoded as a string, array of strings, array of tokens, or array of token arrays. | |
echo |
OptionalNullable[bool] | ➖ | Echo back the prompt in addition to the completion | |
frequency_penalty |
OptionalNullable[float] | ➖ | Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim. | |
max_tokens |
OptionalNullable[int] | ➖ | The maximum number of tokens that can be generated in the completion. | |
presence_penalty |
OptionalNullable[float] | ➖ | Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics. | |
seed |
OptionalNullable[int] | ➖ | If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. | |
stop |
OptionalNullable[models.CreateCompletionStop] | ➖ | Up to 4 sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence. | |
temperature |
OptionalNullable[float] | ➖ | What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. | |
top_p |
OptionalNullable[float] | ➖ | An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. | |
n |
OptionalNullable[int] | ➖ | How many completions to generate for each prompt. Note: Because this parameter generates many completions, it can quickly consume your token quota. | |
user |
Optional[str] | ➖ | A unique identifier representing your end-user, which can help OpenAI to monitor and detect abuse. | |
name |
Optional[str] | ➖ | The name to display on the trace. If not specified, the default system name will be used. | |
fallbacks |
List[models.CreateCompletionFallbacks] | ➖ | Array of fallback models to use if primary model fails | |
retry |
Optional[models.CreateCompletionRetry] | ➖ | Retry configuration for the request | |
cache |
Optional[models.CreateCompletionCache] | ➖ | Cache configuration for the request. | |
load_balancer |
Optional[models.CreateCompletionLoadBalancer] | ➖ | Load balancer configuration for the request. | |
timeout |
Optional[models.CreateCompletionTimeout] | ➖ | Timeout configuration to apply to the request. If the request exceeds the timeout, it will be retried or fallback to the next model if configured. | |
thinking |
OptionalNullable[models.CreateCompletionThinking] | ➖ | Configuration for the thinking mode capability. Set type to adaptive for models that support adaptive thinking (e.g. Claude Opus 4.6, Sonnet 4.6), or enabled with budget_tokens for manual control. |
|
plugins |
List[models.PIIRedactionPlugin] | ➖ | Request-scoped transforms applied to the text exchanged with the model. Currently supports pii_redaction, which replaces PII with placeholders before the provider sees it and restores the original values in the response. |
|
orq |
Optional[models.CreateCompletionOrq] | ➖ | : warning: ** DEPRECATED **: This will be removed in a future release, please migrate away from it as soon as possible. Leverage Orq's intelligent routing capabilities to enhance your AI application with enterprise-grade reliability and observability. Orq provides automatic request management including retries on failures, model fallbacks for high availability, identity-level analytics tracking, conversation threading, and dynamic prompt templating with variable substitution. |
{ "retry": { "count": 3, "on_codes": [ 429, 500, 502 ] }, "fallbacks": [ { "model": "openai/gpt-5" }, { "model": "anthropic/claude-4-opus" } ], "identity": { "id": "identity_01ARZ3NDEKTSV4RRFFQ69G5FAV", "display_name": "Jane Doe", "email": "jane.doe@example.com" }, "thread": { "id": "thread_01ARZ3NDEKTSV4RRFFQ69G5FAV", "tags": [ "customer-support" ] }, "inputs": { "customer_name": "John Smith", "issue_type": "billing" }, "cache": { "ttl": 3600, "type": "exact_match" }, "knowledge_bases": [ { "knowledge_id": "knowledge_01ARZ3NDEKTSV4RRFFQ69G5FAV", "top_k": 5 } ], "timeout": { "call_timeout": 30000 }, "security": { "mask": [ "input", "system" ] } } |
stream |
Optional[bool] | ➖ | N/A | |
retries |
Optional[utils.RetryConfig] | ➖ | Configuration to override the default retry behavior of the client. |
models.CreateCompletionResponse
| Error Type | Status Code | Content Type |
|---|---|---|
| models.APIDefaultError | 4XX, 5XX | */* |