Skip to content

Export sparse model-routing attributes to OpenTelemetry - #67240

Merged
pelikhan merged 18 commits into
mainfrom
copilot/export-model-routing-attributes
Oct 10, 2026
Merged

pelikhan merged 18 commits into
mainfrom
copilot/export-model-routing-attributes

Conversation

Copilot AI commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Export low-cardinality routing metadata once, on the agent conclusion span. Use agent-provenance session records when available, with aw_info.json and proxy-log fallbacks; document the attributes and exclude prompt-derived details.

  • Attributes: Add routing failure code, objective, classification labels, degraded flag, and deviated-request count. Keep existing status, mode, and router-version attributes; omit routing attributes for non-routed runs and non-agent spans.
  • Privacy and scope: Do not export ranked choices, selected IDs, conversation hashes, rationale, degraded-reason text, or per-request details.
  • Classifier cost: Document gh-aw.model_routing.classifier_aic as separate from gh-aw.aic, but defer emission while its attribution prerequisite remains open.
gh-aw.model_routing.status = "selected"
gh-aw.model_routing.objective = "cost"
gh-aw.model_routing.deviated_requests = 1

Copilot AI and others added 2 commits October 9, 2026 17:33
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Copilot AI changed the title [WIP] Export a sparse set of model-routing attributes and document them Export sparse model-routing attributes to OpenTelemetry Oct 9, 2026
Copilot AI requested a review from SivaKesava1 October 9, 2026 17:37
@SivaKesava1
SivaKesava1 marked this pull request as ready for review October 9, 2026 17:53
Copilot AI balanced review requested due to automatic review settings October 9, 2026 17:53

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Classification labels need fixed enum validation to prevent prompt-derived or high-cardinality telemetry values.

1 open finding
What changed in this PR

Exports sparse model-routing metadata on agent conclusion spans and documents the telemetry contract.

Changes:

  • Adds session-first routing summary resolution with metadata and proxy-log fallbacks.
  • Emits allow-listed routing attributes and adds coverage for routing scenarios.
  • Documents routing attributes, scope, exclusions, and deferred classifier cost.
File Description
actions/​setup/​js/​model_attribution.cjs Resolves routing summaries from available sources.
actions/​setup/​js/​model_attribution.test.cjs Tests summary resolution and fallbacks.
actions/​setup/​js/​send_otlp_span.cjs Emits routing attributes on agent conclusion spans.
actions/​setup/​js/​send_otlp_span.test.cjs Tests span scope and emitted attributes.
docs/​src/​content/​docs/​reference/​open-telemetry-attributes.mdx Documents model-routing telemetry.

🧠 Review effort: Balanced

Comment thread actions/setup/js/model_attribution.cjs Outdated
@gh-aw-bot

Copy link
Copy Markdown
Collaborator

@copilot address the following outstanding work in one pass:

  1. Update this branch with the latest main using make merge-main, resolving any conflicts and preserving the intended changes.
  2. Review (actions/setup/js/model_attribution.cjs:241): firstValidated only checks identifier syntax, so an arbitrary classifier output such as private_user_prompt is accepted and exported. Since these labels are derived from prompt classification, this violates the fixed-enum/low-cardinality privacy boundary and can leak prompt-derived text into telemetry. Validate each field against the router's canonical task-type, scope, and complexity allowlists, dropping unknown values. - Export sparse model-routing attributes to OpenTelemetry #67240 (comment)
  3. Fix failing check JS Tests (shard 4/4) (FAILURE): https://github.com/github/gh-aw/actions/runs/37969711932/job/113953109247.

Push the necessary fixes, reply to each listed review thread and resolve it when addressed. Ignore feedback already answered or resolved. Use the pr-finisher skill and stop when only human review or CI remains; do not trigger CI.

Sous-chef head: 3f584e9
Sous-chef work: 31f7e812d45f0792093269b7b90445aa9803225a204ec130694b9ceb0227d7cc 3242b3fb2f699db8dbd37e11a5d94e810de8bac667ebfc9db6a1fd03042724c0
Sous-chef state: c9b0d62763e86d7f92593660dd1c0409fd07fa524d65dbe533cf77f149d4a2dd

Generated by 👨‍🍳 PR Sous Chef · pi · haiku45 · 5.91 AIC · ⌖ 7.1 AIC · ⊞ 1K · ◷
Comment /souschef to run again

Copilot AI and others added 2 commits October 9, 2026 18:39
…routing-attributes

Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>
Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>
Copilot AI requested a review from gh-aw-bot October 9, 2026 19:00
@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot Please remove the work-queue changes from this PR: work_queue_compaction.test.cjs (node:test → vitest import) and work_queue_issues_checks.cjs (new git.getRef/getCommit/getTree/getBlob mocks). They fix tests that are failing on main, which is unrelated to OTEL routing attributes and belongs in its own change. One fix per PR. model_attribution.test.cjs and send_otlp_span.test.cjs pass at bdd3e3b1 on Linux.

Copilot AI and others added 2 commits October 9, 2026 20:23
…routing-attributes

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

@copilot address the following outstanding work in one pass:

  1. Update this branch with the latest main using make merge-main, resolving any conflicts and preserving the intended changes....

Merged latest main in 4380e75; fixed-enum classification validation was added in bdd3e3b, and the unrelated work-queue diffs were removed in 588295f. The routing tests, lint, and final progress validation pass. The reported shard still reproduces unrelated work-queue failures after those changes were removed, so I left those fixes out as requested.

@copilot Please remove the work-queue changes from this PR: work_queue_compaction.test.cjs (node:test → vitest import) and work_queue_issues_checks.cjs (new git.getRef/getCommit/`getTree...

Removed the work_queue_compaction.test.cjs Vitest import and the added Git API mocks from work_queue_issues_checks.cjs in 588295f; both files now match main. The shard still fails on those unrelated baseline tests.

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
@SivaKesava1

Copy link
Copy Markdown
Collaborator

Re-verified 5e40723e: the work-queue files now match main, the PR touches only model_attribution, send_otlp_span, their tests and the OTEL attributes reference, and model_attribution.test.cjs plus send_otlp_span.test.cjs pass on Linux (397 tests).

@copilot /pr-finisher

…routing-attributes

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot #67092 has merged, so the classifier-cost attribution gh-aw.model_routing.classifier_aic depends on is now on main (the unified firewall.token_usage events carry purpose/path/xInitiator). Please merge main and make sure classifier_aic is derived from it (non-zero for a routed run with a routing_classification request), as #67237 describes.

Copilot AI and others added 2 commits October 9, 2026 21:42
…routing-attributes

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

@copilot #67092 has merged, so the classifier-cost attribution gh-aw.model_routing.classifier_aic depends on is now on main (the unified firewall.token_usage events carry purpose/path/`xInit...

Merged current main in f1d9a77. 832c6ed derives classifier_aic from agent-provenance firewall.token_usage records with purpose: routing_classification, emits it as a double on the agent conclusion span, and keeps it separate from gh-aw.aic. Added tests verifying a non-zero routed classifier cost.

@SivaKesava1

Copy link
Copy Markdown
Collaborator

Ready for review. Verified 832c6edb on Linux (model_attribution.test.cjs + send_otlp_span.test.cjs, 398 tests). Routing attributes are emitted once on the agent conclusion span; classifier labels pass through fixed allowlists; classifier_aic comes from #67092's routing_classification attribution and stays separate from gh-aw.aic; the attributes are documented in open-telemetry-attributes.mdx. The unrelated work-queue changes were removed. Review threads resolved.

@SivaKesava1
SivaKesava1 requested a review from pelikhan October 9, 2026 22:40
@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot After the latest main merge (0c0254dc), the cross-module contract test from #67254 fails on this branch:

model_routing_contract.test.cjs › reports the harness selection even when proxy attribution rejects it
expected 'selected' to be 'rejected'   (line 175: result.otel["gh-aw.model_routing.status"])

Root cause: resolveModelRoutingSummary in model_attribution.cjs picks status (and failure_code) from model_routing.outcome first ([outcome?.status, sessionRouting.status, …, infoRouting?.status, …]). That event comes from the agent-writable agent/awf-routing-outcome.json, so the span reports the harness's claim ("selected") even when the runner-written, proxy-corroborated attribution in aw_info.json rejected it. It's the same trust-boundary problem as #67279 (#66963's rule: attribution only from runner-written data).

Please change the precedence:

  • status, failure_code, mode and router_version come from the runner-written routing (aw_info.json model_routing via resolveEffectiveModel, or workflow.info.modelRouting with workflow provenance), never from model_routing.outcome.
  • objective, the classifier labels, degraded, classifier_aic and deviated_requests can come from the proxy-written firewall.model_routing / firewall.token_usage records (firewall provenance), as now.
  • Don't read model_routing.outcome for any exported attribute.

Then make sure model_routing_contract.test.cjs passes together with this PR's tests.

@pelikhan please hold this one until that's fixed.

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

@copilot After the latest main merge (0c0254dc), the cross-module contract test from #67254 fails on this branch:

Updated in a9558df: routing status, failure code, mode, and router version now come from runner-written attribution (aw_info.json, with provenance-checked workflow.info fallback). model_routing.outcome is no longer read for routing telemetry. The cross-module contract and routing/OTEL tests pass.

@SivaKesava1

Copy link
Copy Markdown
Collaborator

@pelikhan Fixed in a9558dfe: routing status, failure code, mode and router version now come only from runner-written attribution (aw_info.json, with a provenance-checked workflow.info fallback), and model_routing.outcome is no longer read for telemetry. Re-verified on Linux: model_attribution, model_attribution_security, send_otlp_span, model_routing_contract (#67254) and unified_session suites pass (573 tests). No open review threads. Ready for review again.

@pelikhan
pelikhan merged commit 178cc11 into main Oct 10, 2026
12 of 13 checks passed
@pelikhan
pelikhan deleted the copilot/export-model-routing-attributes branch October 10, 2026 10:46
@github-actions github-actions Bot mentioned this pull request Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

OpenTelemetry: export a sparse set of model-routing attributes and document them

5 participants