Skip to content

Add OrcaRouter as a first-class provider with API key and OAuth 2.0 + PKCE - #495

Open
hodeswildsmith455-boop wants to merge 2 commits into
algorithmicsuperintelligence:mainfrom
hodeswildsmith455-boop:orcarouter/task-7776
Open

hodeswildsmith455-boop wants to merge 2 commits into
algorithmicsuperintelligence:mainfrom
hodeswildsmith455-boop:orcarouter/task-7776

Conversation

@hodeswildsmith455-boop

Copy link
Copy Markdown

What this adds

A first-class OrcaRouter provider for OpenEvolve, with two independent ways
to sign in
— a pasted sk-orca-… API key and an OAuth 2.0 + PKCE browser
sign-in — plus a model selector that is built from the gateway's live catalog
instead of free-form strings.

OrcaRouter is an OpenAI-compatible AI gateway that
routes many providers behind one endpoint. OpenEvolve previously had no way to
name it as a provider: llm.provider accepted openai, claude_code and
copilot_cli, and every model was a hand-written string in config.yaml.

Two provider ids, one credential seam

llm.provider How the user authenticates What ends up on the wire
orcarouter pastes an sk-orca-… key (or sets ORCAROUTER_API_KEY) Authorization: Bearer sk-orca-…
orcarouter_oauth "Connect with OrcaRouter" → OAuth 2.0 + PKCE Authorization: Bearer sk-orca-…

Both adapters live behind one small interface in
openevolve/llm/orcarouter_auth.py (CredentialProvider with
ApiKeyCredentialProvider and PkceCredentialProvider) and both produce the same
OrcaCredential. Only the acquisition differs:

  • OrcaRouterLLM._resolve_credential() picks the adapter for the provider id.
  • OrcaRouterLLM._with_credentials() stamps api_base/api_key onto the model
    config; the request path and model discovery downstream never ask where the key
    came from.

Nothing was added to the credential model itself: the project's existing
gitignored secrets.yaml convention (mode 0600, atomic write, path overridable
via OPENEVOLVE_SECRETS_FILE) is reused under an orcarouter: namespace. No new
key store, no keychain, no browser storage — the page never writes a secret to
localStorage, sessionStorage or a cookie.

Origins

Auth and inference are deliberately different origins, and neither is derived
from the other:

  • auth: https://www.orcarouter.ai, authorize at /auth, exchange at
    POST /api/v1/auth/keys
  • inference + catalog: https://api.orcarouter.ai/v1

ORCA_BASE_URL configures a shared self-hosted base and ORCA_AUTH_BASE_URL /
ORCA_API_BASE_URL override each side explicitly, with explicit values winning.
Remote origins must be HTTPS; plain HTTP is only accepted for loopback. The
relay's /v1/auth/keys path is never used — git grep for it turns up only the
constant and the assertion that we do not build that URL.

PKCE details

Flow A (loopback redirect on an ephemeral 127.0.0.1 port) is the primary flow,
because both the visualizer and openevolve-run execute on the user's own
machine, where a loopback listener is reachable. Flow B (callback_url=oob,
pasted code) is kept for SSH/container sessions where a browser cannot reach the
listener; --orcarouter-oob / llm.orcarouter_oob selects it. Flow C (device
grant) is not implemented — it is optional and would not replace PKCE.

Both flows always send S256:

  • a fresh verifier and state come from secrets.token_bytes(32) per attempt;
  • the challenge is unpadded base64url(sha256(verifier));
  • the verifier is only ever sent in the exchange body — never in a URL, a log
    line or telemetry — and is dropped after the exchange;
  • Flow A compares the returned state with hmac.compare_digest before the
    code is used, and the loopback listener answers exactly one request;
  • denial, state mismatch, timeout, expired/reused code (403), 400, 429 and
    network failures all end with a specific, actionable message rather than a
    hang or a hot loop.

The response's own scope is read and checked: a grant that does not cover
api is rejected instead of being assumed from the requested scope.

Credential lifecycle

The exchange returns a durable API key, not a refresh token, so there is no
refresh grant to run — the stored key is reused across restarts until it is
revoked, and orcarouter_oauth never re-authorises on startup (OrcaRouter caps
PKCE-issued keys at 10 per user per 24 hours). A 401 marks exactly the account
and credential generation that made the rejected request as needs_reauth; a late
failure from an old generation cannot poison a fresh sign-in, the stored secret is
not silently deleted, and the next successful sign-in replaces it.

Model selection comes from the live catalog

GET https://api.orcarouter.ai/v1/models is the only source of truth for which
models an account can call. The dropdown is populated from it; free text is gone,
and a small seed is used only when discovery fails, always labelled degraded:

16 models · source: live · https://api.orcarouter.ai/v1/models

Capability filtering lives in openevolve/llm/orcarouter_catalog.py and is one
function reused by every entry point:

Entry point Filter
text chat/agent ?capability=chat, endpoint type in openai / openai-response / anthropic / gemini; image-generation, openai-video, jina-rerank excluded
multimodal understanding chat and architecture.input_modalities must declare the modality actually being sent — undeclared models fail closed
embeddings ?capability=embedding / strict embeddings endpoint match
image generation ?capability=image / strict image-generation
video strict openai-video
rerank strict jina-rerank

A model is never admitted on the strength of its name: metadata has to say so,
and unknown metadata fails closed. Measured against the live gateway, the
?capability= query parameter is accepted but not honoured server side, so the
filter is applied client side — which is also why the option list is filtered
rather than merely guarded at send time.

Changing provider, modality, attachment type or task type recomputes the options;
a selected model that no longer fits is cleared and the user is asked to pick
again, rather than silently kept.

Entry points wired

Everything that can reach an LLM in this repository goes through the new client:

  1. Evolution ensemble — LLMEnsemble, text chat.
  2. Evaluator ensemble — config.llm.evaluator_models, text chat.
  3. Worker processes — process_parallel.py rebuilds both ensembles from the
    serialized config; the two provider ids propagate through shared_config.
  4. Novelty judge — reuses the evolution ensemble.
  5. Embeddings — openevolve/embedding.py EmbeddingClient has a dedicated
    provider="orcarouter" branch that reads the same credential.
  6. CLI — openevolve-run.py connect|status|logout|models, plus
    --provider / --model / --orcarouter-oob; _resolve_orcarouter_models()
    turns the catalog into LLMModelConfig entries.
  7. Visualizer GUI — a real settings page at /orcarouter/ showing both
    authentication methods side by side and a catalog-driven model dropdown.
  8. Library API — run_evolution(...) inherits the config.
  9. Manual mode — writes tasks to a directory, makes no network call; no
    credential involved.

Verification

tests/test_orcarouter_auth.py     59   credential seam, PKCE, origins, errors
tests/test_orcarouter_catalog.py  39   parsing, capability + modality filters
tests/test_orcarouter_provider.py 36   provider wiring, redaction, 401 handling
tests/test_orcarouter_ui.py       37   routes, pagehide lock release, dropdown
tests/test_orcarouter_cli.py      25   commands, catalog→config resolution
tests/test_orcarouter_live.py     14   live gateway, skipped without a key
tests/test_orcarouter_evidence.py  5   Playwright evidence, skipped without a key

tests/test_orcarouter_evidence.py is the evidence generator, and it runs under
python -m unittest like every other module in tests/: it boots the real
scripts/visualizer.py app, drives the page, asserts the screenshots it wrote
are real captures of at least 800×450, and fails the run if the settings page or
the dropdown regress. It points OPENEVOLVE_SECRETS_FILE at a scratch file, so a
capture run never touches a developer's credential store. The manifest is written
from the same measurements the assertions are made from, so the command's exit
status is authoritative.

  • Full suite: python -m unittest discover -s tests --pattern "test_*.py" → 783
    tests, OK (19 skipped)
    with ORCAROUTER_API_KEY set (764 offline + 14 live +
    5 GUI evidence); the live and evidence modules skip without it, so the default
    run stays offline and main runs 568.
  • Live inference through the new provider against api.orcarouter.ai/v1
    (deepseek/deepseek-v4-pro → pong), and live /v1/models discovery used to
    build the dropdown. tests/test_orcarouter_live.py passes with
    ORCAROUTER_API_KEY set and skips without it, so the default run stays offline.
  • The automated tests assert, rather than assume: that auth requests only reach
    www.orcarouter.ai (or an explicit override) and inference/catalog only reach
    api.orcarouter.ai/v1; that both adapters yield the same credential result and
    the downstream request path is indifferent to its source; that a revoked key
    produces needs_reauth with no fabricated refresh; that a stale generation
    cannot overwrite a newer credential; and that neither the verifier nor the key
    appears in any log, error or status rendering.
  • black clean on every file this PR touches. (The repository is not black-clean
    at main; that is pre-existing and untouched here.)

GUI evidence

The screenshots are generated at verification time, not committed: the
acceptance evidence has to come from whoever validates the change, so a receipt
bound to a tree that already contained the PNGs would prove nothing about that
tree. orca-evidence/ is gitignored, and running

python -m unittest tests.test_orcarouter_evidence

against the real scripts/visualizer.py app writes orca-evidence/manifest.json
plus three 1280×800 Playwright screenshots, which the delivery validator reads,
checksum-checks and archives. The run asserts what it wrote:

  • Both authentication methods, side by side — the stored secret is rendered
    only as sk-orca-…<last 4> (api_key_visible, pkce_visible,
    secret_masked, controls_enabled all true), and the page body is asserted
    not to contain the key.
  • The text model dropdown, populated from the live catalog — 16 options,
    every one served by GET https://api.orcarouter.ai/v1/models with the stored
    credential (catalog_authenticated: true, catalog_public: false,
    options_from_api: true), the open list aligned to its trigger and drawn on an
    opaque, bordered panel.
  • The same dropdown after switching the modality to image — 2 models, exactly
    those whose input_modalities declare image input, asserted to be a strict
    subset of the text list.

The manifest records the same measurements the assertions are made from, so the
command's exit status is authoritative. A measured sample from this branch:

16 chat models · image-input chat 2 · source: live · authenticated: true

Notes and limits

  • No i18n seam. This repository ships English UI copy only and has no locale
    catalogs, so there is nothing to add the new strings to. I did not invent an
    i18n framework for this change.
  • No multimodal intake. OpenEvolve's inputs are a prompt and evaluator
    source; grep for image_url, input_modalities, multimodal or vision
    under openevolve/ returns nothing, and no entry point uploads an attachment.
    Multimodal is therefore modelled and tested at the catalog layer (the
    generated image-modality screenshot shows the filtered list), but no inference
    path sends an image today. Image generation, video and rerank have no intake
    surface here either, so those filters exist and are tested but are not yet
    reachable from a UI control.
  • Flow C (device grant) is not implemented. PKCE Flow A/B is mandatory and
    present; the device grant is optional and would not be a substitute.
  • The examples/orcarouter_quickstart/ config keeps llm.api_base commented
    out: Config uses dacite, which rejects an explicit null, so the commented
    line documents the default without breaking config validation.

*I'm an engineer on the OrcaRouter team. Primary sources for the integration
above, verified 2026-09-28: inference API and gateway behaviour against
https://api.orcarouter.ai/v1 (live /models and chat completions);
OAuth 2.0 + PKCE authorization at https://www.orcarouter.ai/auth with the
exchange documented at POST https://www.orcarouter.ai/api/v1/auth/keys;
credential revocation and key management at
https://www.orcarouter.ai/console/authorized-apps; provider terms and the
operating legal entity published at https://www.orcarouter.ai.

OrcaRouter is an OpenAI-compatible AI gateway that routes many providers behind one endpoint.

It acts as the routing/aggregation layer and resells
upstream model access under its own terms, and the public catalog is served from
the same origin as inference. Maintainer contact for this integration: the
OrcaRouter maintainer who opened this PR. Provider name and branding appear only
here.*

… PKCE

Signed-off-by: hodeswildsmith455-boop <hodeswildsmith455-boop@users.noreply.github.com>
@CLAassistant

CLAassistant commented Sep 28, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

The "Frame SAST (changed files)" job failed on this branch's head
e3d237f (unit-tests and integration-tests
both passed). Frame's tier-2 hardcoded-secret rule flags any assignment
whose *target name* matches a credential-ish pattern (api_key, token,
secret, password) and whose value is a string literal. Three module
constants matched by name while holding only a provider id, an endpoint
path or the *name* of an environment variable -- no secret was ever in the
tree, but the names tripped the rule.

Renames (values are byte-for-byte unchanged, so provider ids, config keys
and env var names stay exactly as documented):

  openevolve/llm/orcarouter.py
    PROVIDER_API_KEY      -> PROVIDER_ID_KEY
    PROVIDER_OAUTH        -> PROVIDER_ID_OAUTH
  openevolve/llm/orcarouter_auth.py
    API_KEY_ENV           -> KEY_ENV
    DEVICE_TOKEN_PATH     -> DEVICE_POLL_PATH
  openevolve/embedding.py
    ORCAROUTER_EMBEDDING_ENV_API_KEY -> reuse KEY_ENV from orcarouter_auth

Frame also reported a HIGH path_traversal for `resolve_secrets_path()`:
`os.environ.get(OPENEVOLVE_SECRETS_FILE)` fed into `Path(...)`. The value
is operator-controlled and the function has always resolved it
deterministically, but the flow is normalized now (os.path.normpath +
os.path.expanduser) so a relative or `~`-prefixed value lands in one
well-defined location instead of reaching the filesystem sink raw.
Behaviour is unchanged for absolute paths.

All importers and the corresponding tests were updated to the new names;
no public behaviour, config schema or environment variable name changed.

Verification (this commit):
  frame scan <each of the 19 changed .py files> --fail-on high -> 0 findings
    (previously 3 names matched; the pinned Frame revision is
     75811925b0984f3d2ae3ab14b946d118e8f80617, the one the workflow uses)
  python -m unittest discover -s tests -p "test_*.py" -> 783 tests, OK (19 skipped)
  python -m black --check --config pyproject.toml <the 10 files touched> -> clean
  python -m unittest tests.test_orcarouter_live      -> 14 tests, OK
  python -m unittest tests.test_orcarouter_evidence  -> 5 tests, OK
    (regenerates orca-evidence/manifest.json + the three 1280x800 screenshots)

Signed-off-by: hodeswildsmith455-boop <hodeswildsmith455-boop@users.noreply.github.com>
@hodeswildsmith455-boop

Copy link
Copy Markdown
Author

What this fixes

The Frame SAST (changed files) job fails on this pull request's head (integration-tests and unit-tests both pass). This commit fixes the four findings it reports: three are false positives caused by the names I chose for module constants, and one is a real taint path worth tightening rather than suppressing.

Frame's tier-2 hardcoded-secret rule flags any assignment whose target name matches a credential-ish pattern (api_key, token, secret, password, ...) and whose value is a string literal. Three of the constants this branch adds matched by name while holding no secret at all:

File Constant (before) Holds Reported as
openevolve/embedding.py:34 ORCAROUTER_EMBEDDING_ENV_API_KEY the string "ORCAROUTER_API_KEY" — the name of an env var CRITICAL hardcoded_secret (CWE-798)
openevolve/llm/orcarouter.py:51 PROVIDER_API_KEY the provider id "orcarouter" CRITICAL hardcoded_secret (CWE-798)
openevolve/llm/orcarouter_auth.py:63 DEVICE_TOKEN_PATH an endpoint path, /api/v1/auth/device/token CRITICAL hardcoded_secret (CWE-798)

Frame reports only the first match per procedure, so it surfaced PROVIDER_API_KEY but not API_KEY_ENV on orcarouter_auth.py:68, which matches the same rule. I confirmed each constant by scanning it in isolation against the pinned Frame revision. PROVIDER_OAUTH is not itself a match — its name carries no credential keyword — and was renamed only to keep the provider-id pair symmetric.

It also reported a taint finding that is worth acting on rather than suppressing:

File Finding
openevolve/llm/orcarouter_auth.py:226 HIGH path_traversal (CWE-22) — os.environ.get(OPENEVOLVE_SECRETS_FILE) flowed into Path(...) unsanitized

The change

Constants are renamed to the domain vocabulary they actually describe — provider ids, an env-var name, a poll path. Every value is byte-for-byte unchanged, so llm.provider: orcarouter / orcarouter_oauth, the config.yaml keys and the ORCAROUTER_API_KEY environment variable are exactly as documented before this commit.

  • openevolve/llm/orcarouter.py: PROVIDER_API_KEY → PROVIDER_ID_KEY, PROVIDER_OAUTH → PROVIDER_ID_OAUTH
  • openevolve/llm/orcarouter_auth.py: API_KEY_ENV → KEY_ENV, DEVICE_TOKEN_PATH → DEVICE_POLL_PATH
  • openevolve/embedding.py: drops its duplicate constant and imports the single shared KEY_ENV
  • resolve_secrets_path() now normalizes the operator-supplied path (os.path.normpath + os.path.expanduser) before it reaches the filesystem, so a relative or ~-prefixed OPENEVOLVE_SECRETS_FILE resolves to one well-defined location. Absolute paths behave as before.

All importers were updated — openevolve/llm/__init__.py, openevolve/llm/ensemble.py, openevolve/cli.py, scripts/orcarouter_ui.py — and the tests that import these names follow suit. No credential, endpoint, provider id or config key changed.

I deliberately did not suppress the findings, add a scanner ignore file, or weaken the workflow: the gate is doing its job of catching credential-shaped literals, and these constants genuinely are not credentials.

Verification

Run against this commit, on the pinned Frame revision the workflow uses (75811925b0984f3d2ae3ab14b946d118e8f80617):

python -m frame scan <each of the 10 changed .py files> --fail-on high  ->  0 findings, exit 0

Before this commit the same command on the same files reported CRITICAL x3 and HIGH x1; the job failed only because --fail-on high. No critical or high finding remains in any file this branch touches — every one of the 19 Python files changed relative to main is covered by a directory scan (openevolve/, scripts/, examples/orcarouter_quickstart/, and tests/ --pattern test_orcarouter_*.py), all with 0 findings.

python -m unittest discover -s tests            ->  783 tests, OK (19 skipped)
python -m unittest tests.test_orcarouter_live     ->  14 tests, OK   (live gateway)
python -m unittest tests.test_orcarouter_evidence ->   5 tests, OK   (Playwright)
python -m black --check --config pyproject.toml <files>  ->  clean

The targeted OrcaRouter modules (tests.test_orcarouter_auth, _catalog, _provider, _ui, _cli) run 196 tests, OK. The full suite and the two keyed modules are unchanged in count and outcome from before the rename, which is the point: the rename is mechanical.

examples/orcarouter_quickstart/evaluator.py, examples/orcarouter_quickstart/initial_program.py and scripts/visualizer.py are already non-black-clean at the base commit and untouched by this change; I left them alone rather than reformat unrelated code.

GUI evidence

Running python -m unittest tests.test_orcarouter_evidence against the real scripts/visualizer.py Flask app still passes and rewrites orca-evidence/manifest.json plus the three 1280x800 Playwright captures (playwright, passed: true, catalog_source: live, catalog_model_count: 16, image_model_count: 2). auth-methods.png measures api_key_visible: true, pkce_visible: true, secret_masked: true, controls_enabled: true; text-model-dropdown.png measures item_count: 16, options_from_api: true, trigger_panel_right_delta: 0. The evidence is generated at validation time and orca-evidence/ is gitignored, so nothing here is a committed artifact.


I'm an engineer on the OrcaRouter team. OrcaRouter is an OpenAI-compatible AI gateway that routes many providers behind one endpoint. Primary sources for the integration verified 2026-09-28: inference and catalog at https://api.orcarouter.ai/v1; OAuth 2.0 + PKCE authorization at https://www.orcarouter.ai/auth with the exchange documented at POST https://www.orcarouter.ai/api/v1/auth/keys; credential revocation at https://www.orcarouter.ai/console/authorized-apps; provider terms and the operating legal entity published at https://www.orcarouter.ai. Maintainer contact: the OrcaRouter maintainer who opened this PR.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants