Skip to content

Model routing: Claude-engine runs recorded as 'routing rejected (unsupported_endpoint)' although AWF routing succeeded #67012

Description

@SivaKesava1

Summary

Every AWF-routed Claude-engine run is recorded as a failed routing, even when the routed model served every request. The Claude harness correctly uses its own compatible endpoint (behaviour added in #66956), but the post-agent attribution in parse_token_usage.cjs (added in parallel in #66963) still applies a strict endpoint check. As a result, aw_info.json, step outputs, agent_usage.json, gh aw audit and safe-output footers all report routing rejected.

Evidence

Two successful runs of routing-test-claude-routed in githubnext/gh-aw-routing-sandbox. Both used the Claude engine with engine.model-routing and model provider github, were compiled at main 5a8304e, and ran on AWF v0.28.44 (proxy schema model-routing/v0.28.49).

Run AWF selection Claude harness Proxy request records Recorded attribution
37844188877 (t01) claude-sonnet-5 medium, /chat/completions mode=awf-routed endpoint=/v1/messages selected_endpoint=/chat/completions 5/5 routed: as_selected, upstream_endpoint: /v1/messages, deviations: ["endpoint"], all 200 rejected / unsupported_endpoint
37844175739 (t10) claude-opus-5 max, /chat/completions same, with claude-opus-5 max 35/35 as_selected on /v1/messages, deviations: ["endpoint"] rejected / unsupported_endpoint

Findings common to both runs:

  • In /reflect routing_models, the selected model advertises supported_endpoints: ["/v1/messages", "/chat/completions"] with candidate_metadata_complete: true.
  • agent/aw_info.json has "model": "agent" and model_routing: {"status": "rejected", "failure_code": "unsupported_endpoint", "detail": "AWF model routing endpoint /chat/completions is not supported by this engine"}. agent_usage.json contains the same block.
  • The agent job outputs model_routing_status=rejected.
  • The created issues' footers read · claude · routing rejected ·, and their hidden markers read model: agent, routed: rejected (githubnext/gh-aw-routing-sandbox#111, githubnext/gh-aw-routing-sandbox#113).

Control: run 37844201341 uses the pi engine with the same Claude selection and is attributed correctly as routed: opus50 max.

Root cause

#66956 and #66963 were developed in parallel, and the attribution path never adopted the endpoint rule that #66956 introduced.

  • Use engine-compatible endpoints for routed models #66956 added allowEndpointOverride to resolveAWFModelRoutingSelection in actions/setup/js/awf_model_routing.cjs. If AWF selects an endpoint the engine does not speak, a non-Copilot harness may use its own endpoint, provided that /reflect routing_models[].supported_endpoints advertises it for the selected model. claude_harness.cjs (CLAUDE_ROUTING_ENDPOINTS = ["/v1/messages"]) and codex_harness.cjs (CODEX_ROUTING_ENDPOINTS = ["/responses"]) pass the flag.
  • Report AWF-routed model and effort across workflow outputs #66963 added resolveModelRoutingOutcome in actions/setup/js/parse_token_usage.cjs. This function keeps its own MODEL_ROUTING_ENDPOINTS table. For a proxy selection record, it returns unsupported_endpoint whenever routing.endpoint is not in that engine's list, before any metadata check runs. Its later call to resolveAWFModelRoutingSelection also omits allowEndpointOverride.
  • AWF currently selects /chat/completions for every Copilot Claude model, so every routed Claude-engine run fails this check. pi passes only because its list contains all three endpoints. Codex has the same latent mismatch whenever AWF selects /chat/completions for a GPT model (not observed in these runs).
  • The harness outcome cannot correct the result. recordAWFModelRoutingOutcome stores no endpoint, and the advisory block in resolveModelRoutingOutcome can only downgrade a status; a matching selected outcome never confirms one.
  • The tests did not catch it. The Report AWF-routed model and effort across workflow outputs #66963 attribution tests in parse_token_usage.test.cjs cover only pi on /responses, and the Use engine-compatible endpoints for routed models #66956 tests are harness-level.

Impact: for runs where routing worked, the following are all wrong:

  • the aw_info.json model field
  • the empty model and model_effort outputs
  • model_routing_status
  • footers and hidden markers
  • agent_usage.json
  • audit and logs model_routing

The routed model and effort are lost from any analysis that uses these fields.

Proposed fix

1. Use one endpoint policy for the harnesses and attribution. Move the per-engine endpoint lists and the override flag into awf_model_routing.cjs as a single shared policy. claude_harness.cjs, codex_harness.cjs, pi_models_json.cjs and parse_token_usage.cjs should all use it, and MODEL_ROUTING_ENDPOINTS should be removed. The attribution check then matches each harness's endpoint selection:

  • claude and pi use allowEndpointOverride, and so does codex, whose harness already passes it. pi already accepts all three endpoints, so the flag only keeps it consistent.
  • copilot stays strict.

2. Record both the effective endpoint and AWF's selected endpoint.

  • Extend recordAWFModelRoutingOutcome to persist endpoint (the endpoint the harness actually configured) and selected_endpoint (AWF's choice), validated against the known endpoint set. The Claude, Codex and pi harnesses should pass both values from result.selection.
  • Extend recordModelRouting in model_attribution.cjs to write both fields into aw_info.json model_routing. Add SelectedEndpoint to AwInfoModelRouting in pkg/cli/logs_models.go so that audit and logs show both values.

3. Prefer a corroborated harness outcome over AWF's raw endpoint. In resolveModelRoutingOutcome, replace the strict endpoint pre-check with these rules:

  • If AWF's endpoint is in the engine's list, behaviour is unchanged.

  • Otherwise, use the effective endpoint from agent/awf-routing-outcome.json only if all of the following hold:

    • the outcome's status is selected;
    • its wire_model and selected_endpoint match the proxy selection record;
    • its effective endpoint is in the engine's list;
    • the proxy's request-stage records for that model are routed: "as_selected" with upstream_endpoint equal to that effective endpoint, and show no deviations other than endpoint.

    readModelRoutingRecord currently keeps only the last selection or failure record, so it must also return the request records.

  • If there is no harness outcome, which is the case for older harnesses and for these two historical runs, fall back to resolveAWFModelRoutingSelection with allowEndpointOverride on agent/awf-reflect.json. That file is agent-writable, as the existing "does not trust selected routing from sandbox-writable reflect data" test notes, so require the same proxy request corroboration before reporting selected.

  • Keep unsupported_endpoint for the case where the model advertises none of the engine's endpoints, and pass through the detail from resolveAWFModelRoutingSelection. Use a distinct code, such as uncorroborated_endpoint, when a harness or reflect claim is not backed by the proxy's request records.

  • Evaluate the advisory outcome as part of this decision rather than afterwards, so that a matching selected outcome can confirm a selection as well as downgrade one.

The proxy-written model-routing.jsonl remains the authority. The harness outcome only decides which compatible endpoint to verify against the proxy's records.

4. Add regression tests.

In actions/setup/js/parse_token_usage.test.cjs:

  • Positive (Claude, override path).
    • Setup:
      • GH_AW_ENGINE_ID=claude.
      • A proxy selection record for claude-opus-5 with effort max on /chat/completions.
      • request records with routed: "as_selected", upstream_endpoint: "/v1/messages" and deviations: ["endpoint"].
      • Reflect routing_models in which the model advertises /v1/messages with complete metadata.
      • A harness outcome with effective endpoint /v1/messages.
    • Expected:
      • model_routing.status is selected, and model is claude-opus-5.
      • endpoint is /v1/messages and selected_endpoint is /chat/completions.
      • The outputs are model=claude-opus-5, model_effort=max and model_routing_status=selected.
      • The footer label from getEffectiveModelLabel is routed: opus50 max.
  • Positive (fallback path). The same setup without a harness outcome, using reflect and proxy corroboration, still yields selected.
  • Negative (endpoint not advertised). The model's supported_endpoints does not include /v1/messages. Expect rejected / unsupported_endpoint, aw_info.json model remaining agent, and an empty model output.
  • Negative (uncorroborated). The harness outcome claims /v1/messages, but the proxy has no matching as_selected request records on that endpoint. Expect rejected with the uncorroborated code.
  • Copilot stays strict. A copilot-engine selection on an endpoint outside its list is still rejected.

Related tests elsewhere:

  • awf_model_routing.test.cjs: recordAWFModelRoutingOutcome persists valid endpoints and drops invalid ones.
  • claude_harness tests: the recorded outcome contains both endpoints.

Verification

  • Replay the downloaded artifacts from 37844188877 and 37844175739 (model-routing.jsonl and awf-reflect.json) through resolveModelRoutingOutcome with GH_AW_ENGINE_ID=claude. The result should be selected, through the fallback path, because these runs have no endpoint in the harness outcome.
  • A fresh routing-test-claude-routed run should produce the footer routed: sonnet50 medium (t01) or routed: opus50 max (t10), with both endpoints recorded in aw_info.json.

Metadata

Metadata

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions