Skip to content

Latest commit

 

History

History
381 lines (260 loc) · 68.2 KB

File metadata and controls

381 lines (260 loc) · 68.2 KB

Evals

Overview

Available Operations

all

List all evaluators in the workspace.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.evals.all(limit=10)

    # Handle response
    print(res)

Parameters

Parameter Type Required Description
limit Optional[int] ➖ A limit on the number of objects to be returned. Limit can range between 1 and 200, and the default is 10
starting_after Optional[str] ➖ A cursor for use in pagination. starting_after is an object ID that defines your place in the list. For instance, if you make a list request and receive 20 objects, ending with 01JJ1HDHN79XAS7A01WB3HYSDB, your subsequent call can include after=01JJ1HDHN79XAS7A01WB3HYSDB in order to fetch the next page of the list.
ending_before Optional[str] ➖ A cursor for use in pagination. ending_before is an object ID that defines your place in the list. For instance, if you make a list request and receive 20 objects, starting with 01JJ1HDHN79XAS7A01WB3HYSDB, your subsequent call can include before=01JJ1HDHN79XAS7A01WB3HYSDB in order to fetch the previous page of the list.
search Optional[str] ➖ N/A
sort Optional[models.QueryParamSort] ➖ N/A
project_id Optional[str] ➖ N/A
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.GetEvalsResponseBody

Errors

Error Type Status Code Content Type
models.GetEvalsEvalsResponseBody 404 application/json
models.APIDefaultError 4XX, 5XX */*

create

Create a new evaluator in the workspace.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.evals.create(request={
        "code": "<value>",
        "type": "python_eval",
        "path": "Default",
        "description": "",
        "key": "<key>",
    })

    # Handle response
    print(res)

Parameters

Parameter Type Required Description
request models.CreateEvalRequestBody ✔️ The request object to use for the request.
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.CreateEvalResponseBody

Errors

Error Type Status Code Content Type
models.CreateEvalEvalsResponseBody 404 application/json
models.APIDefaultError 4XX, 5XX */*

get

Retrieve a single evaluator by ID with more detail than the list endpoint: full type-specific config, owner, domain_id, metadata, enabled, and output_type.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.evals.get(id="01JMDPA3QW5C1V0NJ1PW34T4E5")

    # Handle response
    print(res)

Parameters

Parameter Type Required Description
id str ✔️ N/A
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.GetEvalResponseBody

Errors

Error Type Status Code Content Type
models.GetEvalEvalsResponseBody 404 application/json
models.APIDefaultError 4XX, 5XX */*

delete

Delete an evaluator by its unique identifier.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    orq.evals.delete(id="<id>")

    # Use the SDK ...

Parameters

Parameter Type Required Description
id str ✔️ N/A
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Errors

Error Type Status Code Content Type
models.DeleteEvalResponseBody 404 application/json
models.DeleteEvalEvalsResponseBody 409 application/json
models.APIDefaultError 4XX, 5XX */*

update

Update an evaluator by ID with the provided fields.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.evals.update(id="<id>", path="Default", project_id="01JMDPA3QW5C1V0NJ1PW34T4E5")

    # Handle response
    print(res)

Parameters

Parameter Type Required Description Example
id str ✔️ N/A
type Optional[str] ➖ Evaluator type. Optional on update — inferred from existing evaluator.
path Optional[str] ➖ Legacy alternative to project_id. Project path. Optional on update — the evaluator keeps its current project when both are omitted. Mutually exclusive with project_id. Default
project_id Optional[str] ➖ Unique identifier of the project that owns the evaluator, as returned by GET /v2/projects. Optional on update — the evaluator keeps its current project when omitted; supplying a different id moves it. Mutually exclusive with path. 01JMDPA3QW5C1V0NJ1PW34T4E5
key Optional[str] ➖ N/A
description Optional[str] ➖ N/A
prompt Optional[str] ➖ N/A
output_type Optional[str] ➖ N/A
categories List[str] ➖ N/A
categorical_labels List[models.UpdateEvalCategoricalLabels] ➖ N/A
dataset_id OptionalNullable[str] ➖ N/A
repetitions Optional[float] ➖ N/A
mode Optional[models.UpdateEvalMode] ➖ N/A
model Optional[str] ➖ N/A
jury Optional[models.UpdateEvalJury] ➖ N/A
schema_ Optional[str] ➖ N/A
url Optional[str] ➖ N/A
method Optional[str] ➖ N/A
headers Dict[str, str] ➖ N/A
payload Dict[str, Any] ➖ N/A
code Optional[str] ➖ N/A
guardrail_config Optional[Any] ➖ N/A
version_increment Optional[models.UpdateEvalVersionIncrement] ➖ N/A
version_description Optional[str] ➖ N/A
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.UpdateEvalResponseBody

Errors

Error Type Status Code Content Type
models.UpdateEvalEvalsResponseBody 404 application/json
models.APIDefaultError 4XX, 5XX */*

list_versions

Returns version history for a specific evaluator.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.evals.list_versions(id="<id>")

    # Handle response
    print(res)

Parameters

Parameter Type Required Description
id str ✔️ N/A
limit Optional[int] ➖ Page size, 1-200. Unset uses the server default (10).
starting_after Optional[str] ➖ N/A
ending_before Optional[str] ➖ N/A
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.ListEvaluatorVersionsResponse

Errors

Error Type Status Code Content Type
models.APIDefaultError 4XX, 5XX */*

get_version

Returns a specific version of an evaluator.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.evals.get_version(id="<id>", version_id="<id>")

    # Handle response
    print(res)

Parameters

Parameter Type Required Description
id str ✔️ N/A
version_id str ✔️ N/A
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.GetEvalVersionResponseBody

Errors

Error Type Status Code Content Type
models.APIDefaultError 4XX, 5XX */*

invoke

Runs an evaluator that already exists in the workspace. Accepts either a conversation or the structured input and output fields; when both are present the conversation wins.

Example Usage

from orq_ai_sdk import Orq
import os


with Orq(
    api_key=os.getenv("ORQ_API_KEY", ""),
) as orq:

    res = orq.evals.invoke(id="<id>")

    # Handle response
    print(res)

Parameters

Parameter Type Required Description
id str ✔️ Accepts a bare id, id@version, or id@environment.
context Optional[models.EvaluationContext] ➖ The data to grade. When messages is present it is the conversation and
input.user_query is ignored; output.response is appended only when the
conversation carries no assistant turn.
model Optional[str] ➖ Model to grade with, as a catalog id such as "openai/gpt-4o".

Only meaningful for a hub template of type llm_eval or ragas, which has no
model of its own. A stored evaluator uses the model on its own definition
and ignores this.
query Optional[str] ➖ Latest user message. Folds into context.input.user_query.
output Optional[str] ➖ The generated response from the model. Folds into
context.output.response.
reference Optional[str] ➖ The reference used to compare the output. Folds into
context.input.expected_output.
retrievals List[str] ➖ Knowledge base retrievals. Folds into context.input.retrievals.
messages List[Dict[str, Any]] ➖ The conversation that produced the output. Folds into
context.messages.
variables Dict[str, Any] ➖ Template variables for evaluator prompt substitution. Folds into
context.variables.
retries Optional[utils.RetryConfig] ➖ Configuration to override the default retry behavior of the client.

Response

models.EvaluationResult

Errors

Error Type Status Code Content Type
models.APIDefaultError 4XX, 5XX */*