Repository navigation
Report sub-agent failures, resolved models, and per-agent spend - #66961
Conversation
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
There was a problem hiding this comment.
🟡 Changes recommended
Outcome aggregation, model and credit attribution, legacy pi compatibility, and an outdated integration assertion remain unresolved.
12 open findings
Update actual-model assertion for enriched usage rows · New Emit top-level timestamps for usage records · New Preserve mixed effort across merged usage records · New Resolve aliases using per-agent metric evidence first · New Exclude routing-classifier credits from endpoint reconciliation · New Report main-agent attribution when all sub-agents fail · New Initialize incomplete count for Copilot agent usage · New Support legacy dispatch records with session-scoped correlation · New Initialize incomplete count for Pi agent usage · New Skip usage for zero-token failed or aborted messages · New Aggregate usage under the selected model key · New Derive group model and failure reason order-independently · New
What changed in this PR
Extends gh aw audit and logs reporting with structured sub-agent outcomes, shared model identity resolution, and per-agent usage and spend, addressing #66960.
Changes:
- Tracks failures, incomplete invocations, effort, and served models.
- Adds pi telemetry correlation and per-agent credit attribution.
- Carries attribution through reports, comparisons, schemas, and tests.
| File | Description |
|---|---|
| schemas/logs.schema.json | Adds attribution and cost fields. |
| schemas/logs-jsonl.schema.json | Extends JSONL attribution schemas. |
| schemas/audit.schema.json | Extends audit and comparison schemas. |
| pkg/cli/token_usage_types.go | Defines outcome and usage data. |
| pkg/cli/token_usage_test.go | Updates fallback attribution expectations. |
| pkg/cli/token_usage_subagent.go | Resolves models and attributes credits. |
| pkg/cli/token_usage_subagent_session.go | Parses lifecycle and usage evidence. |
| pkg/cli/token_usage_subagent_session_test.go | Covers failures and agent metrics. |
| pkg/cli/token_usage_parse.go | Records endpoint context. |
| pkg/cli/token_usage_declared_subagents.go | Matches declarations and reports failures. |
| pkg/cli/token_usage_declared_subagents_test.go | Covers aliases and pi attribution. |
| pkg/cli/session_parser_test.go | Updates integration expectations. |
| pkg/cli/model_routing.go | Adds per-agent routing costs. |
| pkg/cli/model_routing_test.go | Tests agent cost aggregation. |
| pkg/cli/model_identity.go | Centralizes model identity resolution. |
| pkg/cli/logs_report.go | Exposes aggregate agent costs. |
| pkg/cli/audit_report_render.go | Renders outcomes and costs. |
| pkg/cli/audit_finding_codes.go | Adds the sub-agent failure code. |
| pkg/cli/audit_cross_run_render.go | Renders cross-run agent spend. |
| pkg/cli/audit_comparison.go | Includes agent costs in comparisons. |
| actions/setup/js/pi_subagent_extension.test.cjs | Tests invocation and result telemetry. |
| actions/setup/js/pi_subagent_extension.cjs | Emits invocation IDs and terminal results. |
| actions/setup/js/pi_session.test.cjs | Tests telemetry transformation. |
| actions/setup/js/pi_session.cjs | Preserves invocation and outcome fields. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
|
@copilot address the following outstanding work in one pass:
Push the necessary fixes, reply to each listed review thread and resolve it when addressed. Ignore feedback already answered or resolved. Use the pr-finisher skill and stop when only human review or CI remains; do not trigger CI. Sous-chef head: 7670436
|
Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>
…-agent-rows Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>
Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>
|
@copilot Tested at
Please add fixtures for 37733652548 (legacy pi) and a pi fixture with this PR's new invocation ids, matching the acceptance table in #66960. I'll re-run the audits on the next head. |
…-agent-rows Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
|
@copilot Re-tested at
Please add the five runs' expected values (from the #66960 table plus the new pi run) as fixtures, so a regression in one engine shows up in tests. I'll re-run the audits on the next head. |
Implemented in |
|
@copilot Re-tested Working:
Regression: run 37505153056 (Copilot, from before the structured sub-agent events) now reports no sub-agents at all ( Minor: for pi, the |
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Metrics-only unified |
|
@copilot Re-tested Fixed: pi Still open, two items:
The |
…-agent-rows Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Merged latest |
|
Re-tested
One minor issue, only on old artifacts: on 37505153056 each sub-agent shows The new warnings ( @copilot /pr-finisher |
|
@copilot Scope update: please don't change the |
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Fixed in |
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Restored the prior instance-count behavior in |
|
Ready for review. Sandbox-verified at |


gh aw auditcould report failed sub-agents as successful, disagree on aliased or dated model IDs, and omit per-agent usage and cost. This change adds structured sub-agent outcomes, shared model identity resolution, and per-agent attribution across audit and logs reports.agent-stdio.logwhen structured events are available.Example attribution:
{ "requested_model": "small", "resolved_model": "gpt-5.4-mini", "served_models": ["gpt-5.4-mini-2026-03-17"], "completed_count": 2, "failed_count": 0 }