Skip to content

Report sub-agent failures, resolved models, and per-agent spend - #66961

Merged
pelikhan merged 18 commits into
mainfrom
copilot/audit-fix-sub-agent-rows
Oct 9, 2026
Merged

pelikhan merged 18 commits into
mainfrom
copilot/audit-fix-sub-agent-rows

Conversation

Copilot AI commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

gh aw audit could report failed sub-agents as successful, disagree on aliased or dated model IDs, and omit per-agent usage and cost. This change adds structured sub-agent outcomes, shared model identity resolution, and per-agent attribution across audit and logs reports.

  • Outcomes and findings: Track completed, failed, and incomplete invocations, effort, and sanitized errors. Failed invocations without served requests no longer receive an effective model; audit reports a “Sub-agent Failed” finding.
  • Model identity: Resolve aliases, provider-qualified patterns, Anthropic version spellings, and dated IDs consistently. Parse pi dispatch, usage, and terminal-result events instead of relying on agent-stdio.log when structured events are available.
  • Usage and cost: Include per-agent requests, tokens, credits, models, effort, and outcomes. Attribute pi credits by matching proxy entries on model, token tuple, and time order; warn on unmatched usage or endpoint-specific credit differences. Preserve existing routing cost buckets while adding main-agent and per-sub-agent splits.
  • Reports and schemas: Carry the breakdown through audit, logs, cross-run summaries, comparisons, console rendering, and the audit/logs JSON schemas.

Example attribution:

{
  "requested_model": "small",
  "resolved_model": "gpt-5.4-mini",
  "served_models": ["gpt-5.4-mini-2026-03-17"],
  "completed_count": 2,
  "failed_count": 0
}

Copilot AI and others added 2 commits October 8, 2026 18:02
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Copilot AI changed the title [WIP] Fix sub-agent rows to correctly report failures and resolve aliases Report sub-agent failures, resolved models, and per-agent spend Oct 8, 2026
Copilot AI requested a review from SivaKesava1 October 8, 2026 18:32
@SivaKesava1
SivaKesava1 marked this pull request as ready for review October 8, 2026 18:33
Copilot AI balanced review requested due to automatic review settings October 8, 2026 18:33

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Outcome aggregation, model and credit attribution, legacy pi compatibility, and an outdated integration assertion remain unresolved.

12 open findings
What changed in this PR

Extends gh aw audit and logs reporting with structured sub-agent outcomes, shared model identity resolution, and per-agent usage and spend, addressing #66960.

Changes:

  • Tracks failures, incomplete invocations, effort, and served models.
  • Adds pi telemetry correlation and per-agent credit attribution.
  • Carries attribution through reports, comparisons, schemas, and tests.
File Description
schemas/​logs.schema.json Adds attribution and cost fields.
schemas/​logs-jsonl.schema.json Extends JSONL attribution schemas.
schemas/​audit.schema.json Extends audit and comparison schemas.
pkg/​cli/​token_usage_types.go Defines outcome and usage data.
pkg/​cli/​token_usage_test.go Updates fallback attribution expectations.
pkg/​cli/​token_usage_subagent.go Resolves models and attributes credits.
pkg/​cli/​token_usage_subagent_session.go Parses lifecycle and usage evidence.
pkg/​cli/​token_usage_subagent_session_test.go Covers failures and agent metrics.
pkg/​cli/​token_usage_parse.go Records endpoint context.
pkg/​cli/​token_usage_declared_subagents.go Matches declarations and reports failures.
pkg/​cli/​token_usage_declared_subagents_test.go Covers aliases and pi attribution.
pkg/​cli/​session_parser_test.go Updates integration expectations.
pkg/​cli/​model_routing.go Adds per-agent routing costs.
pkg/​cli/​model_routing_test.go Tests agent cost aggregation.
pkg/​cli/​model_identity.go Centralizes model identity resolution.
pkg/​cli/​logs_report.go Exposes aggregate agent costs.
pkg/​cli/​audit_report_render.go Renders outcomes and costs.
pkg/​cli/​audit_finding_codes.go Adds the sub-agent failure code.
pkg/​cli/​audit_cross_run_render.go Renders cross-run agent spend.
pkg/​cli/​audit_comparison.go Includes agent costs in comparisons.
actions/​setup/​js/​pi_subagent_extension.test.cjs Tests invocation and result telemetry.
actions/​setup/​js/​pi_subagent_extension.cjs Emits invocation IDs and terminal results.
actions/​setup/​js/​pi_session.test.cjs Tests telemetry transformation.
actions/​setup/​js/​pi_session.cjs Preserves invocation and outcome fields.

🧠 Review effort: Balanced


💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

Comment thread pkg/cli/session_parser_test.go Outdated
Comment thread actions/setup/js/pi_subagent_extension.cjs Outdated
Comment thread pkg/cli/token_usage_declared_subagents.go Outdated
Comment thread pkg/cli/token_usage_subagent.go Outdated
Comment thread pkg/cli/token_usage_subagent.go Outdated
Comment thread pkg/cli/token_usage_subagent_session.go Outdated
Comment thread pkg/cli/token_usage_subagent_session.go
Comment thread pkg/cli/token_usage_subagent_session.go
Comment thread pkg/cli/token_usage_subagent_session.go Outdated
Comment thread pkg/cli/token_usage_subagent_session.go Outdated
@gh-aw-bot

Copy link
Copy Markdown
Collaborator

@copilot address the following outstanding work in one pass:

  1. Update this branch with the latest main using make merge-main, resolving any conflicts and preserving the intended changes.
  2. Review (pkg/cli/session_parser_test.go:226): The exact actual-model assertion at line 226 still expects a request-only row. The new parser also returns resolved/served model fields and this fixture's 6,000 input tokens, 600 output tokens, 60 cache-read tokens, and 120 cache-write tokens. That assertion no longer matches for any of the three source variants. Update it alongside these request-row expectations to check the enriched actual row. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  3. Review (actions/setup/js/pi_subagent_extension.cjs:110): This wrapper has no timestamp, and the transformer does not promote the nested message timestamp. The Go parser therefore gives these usages zero timestamps and sorts them by agent name. If z-reader responds before a-reader, the forward-only proxy matcher handles a-reader first and skips z-reader's earlier entry, losing its credits and preventing the main-agent breakdown. Add a top-level observation timestamp to each emitted usage record. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  4. Review (pkg/cli/token_usage_declared_subagents.go:102): Returning an empty string for conflicting efforts makes a conflict indistinguishable from an uninitialized accumulator. agentRows() merges both usage and request effort: merging low, then high, then that request's high reports high, hiding the mixed effort. Map iteration can change the result. Preserve a sticky mixed value instead of resetting the accumulator. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  5. Review (pkg/cli/token_usage_subagent.go:41): Run-wide candidates are searched before each agent's metric evidence when resolving aliases. With the default small mapping, a run containing haiku traffic and a worker whose metrics split small requests from gpt-5.4-mini credits can resolve that worker to the alphabetically earlier haiku model. Its request and credit metrics then remain assigned to different models. Resolve aliases using configured/per-instance evidence before aggregation; use run-wide IDs only to corroborate model spellings, not to choose another agent's model. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  6. Review (pkg/cli/token_usage_subagent.go:262): TotalAIC includes routing-classifier credits, while main/sub-agent metrics exclude them; the pi main-agent calculation explicitly skips routing_classification entries. Correct attribution therefore triggers a mismatch warning whenever classifier spend exceeds 0.01. Reconcile against the proxy's non-classifier total, preserving the existing run-wide total, so normal routing costs are not reported as endpoint-specific discrepancies. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  7. Review (pkg/cli/token_usage_subagent.go:293): When all pi sub-agents fail before serving requests, requestUsages is empty and this returns before constructing the main-agent row. Proxy records can still contain the main agent's requests and credits, but the report omits them and leaves MainAgentCost at zero. Handle confirmed zero-sub-agent-usage failures separately so main-agent attribution does not require a successful child request. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  8. Review (pkg/cli/token_usage_subagent_session.go:439): The Copilot request row starts as incomplete, but its agent-usage row does not. An invocation without a terminal event therefore reports incomplete_count: 1 in model requests but zero in agent_usage, routing costs, and the console outcome table. Initialize the usage count here; the existing completed/failed handlers already decrement it. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  9. Review (pkg/cli/token_usage_subagent_session.go:476): Pi records emitted before this change have neither invocationId nor agentId; the previous extension and transformer emitted only the agent name. Rejecting those dispatch records aborts structured parsing and falls back to stdio, so existing runs still lose per-agent usage and spend. Add conservative, session-scoped legacy correlation across dispatch and usage records, retaining unknown/incomplete outcomes when terminal evidence is unavailable. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  10. Review (pkg/cli/token_usage_subagent_session.go:494): Pi dispatch starts an incomplete request row but leaves the agent-usage incomplete count at zero. If the child or parent stops before a terminal result, usage and routing-cost reports show an instance with no outcome. Initialize this count on dispatch, matching the request row and the result handler that clears it. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  11. Review (pkg/cli/token_usage_subagent_session.go:535): A pi message_end with a model, stopReason: "error", and no usage still increments served requests here. This is the connection-error shape in pi_provider.test.cjs:316–325. Even after a failed terminal result, rows() assigns an effective model and omits SUBAGENT_FAILED despite no served inference. Preserve the error but skip usage accumulation for zero-token error/aborted messages; keep usage from partially failed responses that consumed tokens. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  12. Review (pkg/cli/token_usage_subagent_session.go:543): When responseModel differs from message.model (for example, a dated served ID), this writes under a different key from the lookup at line 536. Subsequent responses overwrite rather than accumulate, so subagent_model_actuals undercounts requests and tokens while agent_usage retains them. actualRows() also uses the wrong model ID. Store the aggregate under the selected model key. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  13. Review (pkg/cli/token_usage_subagent_session.go:665): Mixed failed/completed invocations produce different reports depending on Go map iteration order. If an unserved failure is visited first, its empty effective model is never replaced by a later completed invocation's model, and its SUBAGENT_FAILED reason survives. The reverse order retains the served model. Collect served-model evidence across the whole group, then derive the effective model and reason once, distinguishing no evidence from conflicting models. - Report sub-agent failures, resolved models, and per-agent spend #66961 (comment)
  14. Fix failing check impacted-go-tests (FAILURE): https://github.com/github/gh-aw/actions/runs/37825190096/job/113476346881.
  15. Fix failing check lint-go-golangci (2) (FAILURE): https://github.com/github/gh-aw/actions/runs/37825190096/job/113476347108.

Push the necessary fixes, reply to each listed review thread and resolve it when addressed. Ignore feedback already answered or resolved. Use the pr-finisher skill and stop when only human review or CI remains; do not trigger CI.

Sous-chef head: 7670436
Sous-chef work: 00eb2d9039d35edbc82950e06a4953b70d5a2ca668d64b0d3f941a9f0608ce8f 20c5fb9e2e37672daf5b7271dae9c95ff59f03cf41d0cde5ac7e2acc9456d2d1 42e5d572d9d75b7168fe7d0aeda0717cc07e65ed240f6d299e6fec661214f1ac 4ad7576d59a0945703860c0a1b99c05d2ff6b411f683dfe4a970c3256fa28ee2 57c6a1efb931eb17302fc32510d99ac6beb7024541dd1e77b53a549d3c01ef0c 80eb32866da791f39e93b6cf59700ed07c6382fcbf2d2c6405a25d831eaacfb3 849639dbd709812a9244eb15ca60b5cd9d2b945797f658218358f3e2867bbdb4 97de7c7b53edcbf37f3ea441b9fe962f2de837ca2c527583e4d74e4d11dd70a9 bf9db078a1debd533ac179a75b05199fdfa39b13ab931758fe67dba0fbac078a cc121bddb0907d343caeb11da6acf4a10fbf0dd180fc75ffd06390b6b5b59d5c cf39595ff93544fe3c44ea7fa4e69641c831ae48df8e29bc0f352ff6fbfead52 d542397180f16d8dbcee60c5bf66f65de56b358026ec3032afd5fbd3680acd61 e5851eba6f000ad390a4003b1d8c4886c2fc5eedf2a0cb7170824599d163ee8d fa6566aa88b0b9ca380f2c1e3fc92258be9ffa5951665af3dd7c16f69883c5d3
Sous-chef state: 99c1194cb1fc15e9643f04203b253e2d759498139172b7aa6f976045cc34293b

Generated by 👨‍🍳 PR Sous Chef · pi · haiku45 · 7.31 AIC · ⌖ 9.08 AIC · ⊞ 1K · ◷
Comment /souschef to run again

Copilot AI and others added 3 commits October 8, 2026 19:15
Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>
…-agent-rows

Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>
Co-authored-by: gh-aw-bot <259018956+gh-aw-bot@users.noreply.github.com>
Copilot AI requested a review from gh-aw-bot October 8, 2026 19:48
@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot Tested at bc0ca7fc with gh aw audit --json on the four runs from #66960, plus a new pi run compiled at this PR's head (t11, two declared sub-agents). What works: failed sub-agents (SUBAGENT_FAILED, error text, "Sub-agent Failed" finding), alias resolution (small → gpt-5.4-mini), and per-agent Copilot costs (main 389.122, research 494.207, explore 40.977 + 9.141, with effort). Four problems remain:

  1. False credit-mismatch warnings. Every run warns that per-agent credits differ from the proxy total, because the proxy total includes the router's classifier request (purpose: routing_classification), and the warning names the classifier's endpoint:
    • 37813733289: per-agent AI credits (17.018) differ from proxy total (17.341) for endpoint /chat/completions. 17.018 is exactly the non-classifier proxy total, so this should not warn.
    • 37505153056: (933.446) differ from proxy total (933.486) for endpoint /responses. The agent traffic was all /v1/messages; /responses and the 0.04 difference are the classifier's.
    • 37813720288: 2.743 vs 2.635 /responses. The real non-classifier gap is 2.743 vs 2.576, from the cache-write accounting difference described in audit: sub-agent rows hide failures, ignore model aliases and dated ids, and have no per-agent spend (follow-up to #66740) #66960.
      Exclude classifier traffic from the comparison, compare per endpoint actually used by agent traffic, and only warn on a real difference.
  2. Legacy pi runs fall back to the heuristic. For 37733652548 (before this PR): invalid pi.subagent_dispatch on line 169: missing pi sub-agent invocation identity, then the stdio fallback and its warning, and agent_usage is null. audit: sub-agent rows hide failures, ignore model aliases and dated ids, and have no per-agent spend (follow-up to #66740) #66960 used this run as an acceptance case (no stdio warning, per-agent credits 1.001 / 1.393). Please accept dispatch/event records without an invocation id by deriving one (for example agent name plus dispatch order) instead of rejecting the source.
  3. New pi run 37832734557 (this PR's extension) is incomplete. The proxy served claude-haiku-4-5-20251001 ×4 (2.426 credits) and gpt-5.6-luna ×3 (0.32). Audit reports file-summarizer with 2 requests and 1.406 credits, but quick-checker with 2 requests and no credits (could not match pi sub-agent quick-checker model claude-haiku-4-5-20251001 usage to proxy token usage), subagent_model_actuals claude-haiku-4.5=2 (should be 4), and no main row in agent_usage. The mismatch warning then compares 1.406 with 2.745 (2.745 includes the main agent).
  4. Minor: in 37813720288, quick-checker's served_models lists the alias small alongside gpt-5.4-mini and gpt-5.4-mini-2026-03-17. Only served wire ids should appear there.

Please add fixtures for 37733652548 (legacy pi) and a pi fixture with this PR's new invocation ids, matching the acceptance table in #66960. I'll re-run the audits on the next head.

Copilot AI and others added 2 commits October 8, 2026 19:55
…-agent-rows

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot Re-tested at 35839302 (same five runs, gh aw audit --json). The legacy pi run is fixed: 37733652548 now shows per-agent credits 1.393 / 1.001, main 0.33, and no warnings. But this head regresses every Copilot run:

  1. The Copilot session parser rejects the whole source. All three Copilot runs warn:
    failed to parse unified subagent information: usage/aw_session.jsonl: invalid tool.execution_complete on line 97: json: cannot unmarshal object into Go struct field sessionSubagentData.error of type string
    agent-session.jsonl: invalid tool.execution_complete on line 35: json: cannot unmarshal object into Go struct field sessionSubagentData.error of type string
    
    In Copilot CLI 1.0.90 events, data.error can be an object (for example on tool.execution_complete). One event that doesn't fit the struct now discards all sub-agent data, so agent_usage is empty, and the failed-sub-agent rows, SUBAGENT_FAILED, and per-agent costs that were correct at bc0ca7fc are gone. Please accept both string and object errors (keep the message text). Ignore event types the sub-agent parser doesn't need instead of failing the source. Add a fixture line with an object error on tool.execution_complete.
  2. The stdio fallback now produces fake sub-agents again. With the structured source rejected, subagent_model_requests contains rows built from tool-call lines such as Check repository status (shell). 37505153056 has 28 of them (Read OpenAPI spec head shell, List all router files shell, …), all REQUESTED_MODEL_NOT_OBSERVED, next to the real Research. 37813720288 has Check repository status shell, and File-summarizer is no longer SUBAGENT_FAILED. This is the false-positive class from audit: sub-agent attribution is guessed from agent-stdio.log, so real sub-agents and their spend are invisible #66740. The fallback should accept only real sub-agent dispatch lines (● Name (model: X) or the legacy Name(model) form), never (shell) or other tool labels, and agent names should keep their original case.
  3. New pi run 37832734557 is still inconsistent. Credits are now inferred (file-summarizer 0.619, quick-checker 0.799), but the audit still warns could not match pi sub-agent … usage to proxy token usage; credits were not inferred for both. The per-agent sum (0.619 + 0.799 + main 0.32 = 1.738) doesn't match the proxy, which served claude-haiku-4-5-20251001 ×4 = 2.426 plus gpt-5.6-luna ×3 = 0.32, so 2.745 total. The two sub-agents' Haiku credits should add up to 2.426, and the warning shouldn't appear when matching succeeds.

Please add the five runs' expected values (from the #66960 table plus the new pi run) as fixtures, so a regression in one engine shows up in tests. I'll re-run the audits on the next head.

Copilot AI commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@copilot Maintainer guidance from @pelikhan: all the information audit/logs need must land in the agent unified session file (usage/aw_session.jsonl, see `docs/src/content/docs/specs/unified-agent-s...

Implemented in 14ccbbd: pi lifecycle, request, model, and outcome data now enters aw_session.jsonl; per-agent pi credits are matched from unified firewall usage; agent-session.jsonl is limited to legacy runs without a usage artifact, and stdio inference is removed. Added Copilot/pi unified-session fixtures and updated the runtime vocabulary. The final change-scoped gate passed.

@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot Re-tested 14ccbbde (gh aw audit --json on the same five runs, plus a fresh pi run compiled at this head).

Working:

  • Fresh pi run 37855181214 (t11, two declared sub-agents): aw_session.jsonl now has subagent.configured/started/request/completed, and the per-agent credits match the proxy total exactly (file-summarizer 1.427 + quick-checker 1.027 + main 0.337 = 2.791).
  • Legacy pi runs 37733652548 and 37832734557: same splits, consistent totals, no warnings.
  • Copilot runs 37813720288 and 37813733289: failed sub-agents counted (3 and 7), small resolves to gpt-5.4-mini, dated ids served correctly, and no fake rows.

Regression: run 37505153056 (Copilot, from before the structured sub-agent events) now reports no sub-agents at all (agent_usage, subagent_model_requests and actuals are all empty), although its usage/aw_session.jsonl session.shutdown.data.agentMetrics has main (389.12 credits), research/subagent-research (494.21), explore/awf-routing (40.98) and explore/ghaw-issues (9.14). The earlier heads reported these through the agent-session.jsonl fallback, which now applies only when the usage artifact is missing. Since agentMetrics is already in the unified file, please build per-agent rows from it even when no valid subagent.* lifecycle events are present (instances and outcome unknown, credits and models from agentMetrics), rather than re-enabling the native-file fallback.

Minor: for pi, the main row has an empty served_models on all three pi runs, although the proxy served the main model (for example gpt-5.6-luna) for those requests.

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@copilot Re-tested 14ccbbde (gh aw audit --json on the same five runs, plus a fresh pi run compiled at this head).

Working:...

Metrics-only unified agentMetrics now produces per-agent model/usage rows without inferring invocation outcomes, and pi main rows include served model identities. Added regression coverage; fixed in commit pending hash verification.

@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot Re-tested 2d429088 (gh aw audit --json, fresh download folder) on the five earlier runs plus the fresh pi run 37855181214.

Fixed: pi main rows now show served_models: gpt-5.6-luna on all three pi runs. Every other row is unchanged and correct (Copilot failed sub-agents 3 and 7, small → gpt-5.4-mini, pi credits adding up to the proxy totals).

Still open, two items:

  1. Run 37505153056 still has no per-agent rows. The metrics-only path never runs, because the whole unified file is rejected first. firewall_token_usage.warnings has: failed to parse unified subagent information: usage/aw_session.jsonl: invalid subagent.configured on line 222: missing agent or tool call identity. That run predates fix: prefer unified sessions for subagent attribution #66746, so its subagent.configured events lack identities, but the same file's session.shutdown.data.agentMetrics is valid (main 389.12, subagent-research 494.21, awf-routing 40.98, ghaw-issues 9.14 credits). Please make parseSessionSubagentModelsDetailed skip a lifecycle event that has no identity (recording a warning) instead of failing the whole file, so the agentMetrics rows still come through. The new regression test should use a unified fixture that contains such an identity-less subagent.configured and valid agentMetrics.

  2. verify-integration-build fails (also on 14ccbbde): pkg/cli/token_usage_subagent_session_test.go:479:53: undefined: tokenUsageEntryModels. That file has no build tag, but it uses helpers defined only in //go:build !integration files (tokenUsageEntryModels in token_usage_declared_subagents_test.go, and copilotSubagentIntegrationMetrics). Reproduce with go vet -tags integration ./pkg/cli.

The JS Tests (shard 2/4) failure (work_queue_documentation.test.cjs) doesn't involve files this PR touches.

Copilot AI and others added 2 commits October 8, 2026 23:30
…-agent-rows

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>
Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

@copilot Re-tested 2d429088 (gh aw audit --json, fresh download folder) on the five earlier runs plus the fresh pi run 37855181214.

Fixed: pi main rows now show served_models: gpt-5.6-luna on...

Merged latest main and fixed both items in 07b8af60: identity-less lifecycle records now emit a diagnostic and are skipped without discarding agentMetrics, and the shared test fixture/helpers compile with the integration build tag. Targeted parser tests and go vet -tags integration ./pkg/cli pass.

@SivaKesava1

Copy link
Copy Markdown
Collaborator

Re-tested 07b8af60 (gh aw audit --json, fresh download folder, same six runs). go vet -tags integration ./pkg/cli passes locally, and verify-integration-build is green.

  • 37505153056 is fixed: main 389.12, subagent-research 494.21, awf-routing 40.98, ghaw-issues 9.14 credits, which match agentMetrics exactly, with served model claude-opus-5.
  • Copilot sub-agent runs 37813720288 and 37813733289, legacy pi runs 37733652548 and 37832734557, and fresh pi run 37855181214 are unchanged and correct. Failed sub-agents are counted (3 and 7), small resolves to gpt-5.4-mini, and the pi credits add up to the proxy totals.
  • JS Tests (shard 2/4) still fails only in work_queue_documentation.test.cjs, which this PR doesn't touch.

One minor issue, only on old artifacts: on 37505153056 each sub-agent shows instance_count: 2, completed_count: 1, although each ran once. Its lifecycle events are keyed by toolCallId (toolu_…), but agentMetrics is keyed by agent UUID, so the two rows merge by name and each adds an instance. A row built only from agentMetrics shouldn't add to instance_count when a lifecycle row for the same agent name exists.

The new warnings (per-agent AI credits … differ from non-classifier proxy total, 2–4% on Copilot runs) look correct to us: they compare Copilot's own agentMetrics accounting with the proxy's.

@copilot /pr-finisher

@SivaKesava1

Copy link
Copy Markdown
Collaborator

@copilot Scope update: please don't change the instance_count behaviour from my previous comment in this PR. It only affects artifacts from before #66746, and we want to land this PR as it is. Limit /pr-finisher to the open review threads, formatting, lint and the impacted tests. Any further gaps will go into separate issues.

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Re-tested 07b8af60 (gh aw audit --json, fresh download folder, same six runs). go vet -tags integration ./pkg/cli passes locally, and verify-integration-build is green.

  • **37505153056 is fix...

Fixed in 22e88624: lifecycle-backed instance counts are authoritative when merged with metrics-only rows of the same agent name, while metrics and spend remain included. Added regression coverage.

Co-authored-by: SivaKesava1 <11771739+SivaKesava1@users.noreply.github.com>

Copilot AI commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

@copilot Scope update: please don't change the instance_count behaviour from my previous comment in this PR. It only affects artifacts from before #66746, and we want to land this PR as it is. L...

Restored the prior instance-count behavior in fd35d48 and limited this pass to formatting, lint, and impacted attribution tests.

@SivaKesava1

Copy link
Copy Markdown
Collaborator

Ready for review. Sandbox-verified at 07b8af60 (gh aw audit --json on six runs: Copilot runs with failed and aliased sub-agents, legacy and fresh pi runs, and a pre-#66746 Copilot run); fd35d482 has no further code changes. All review threads are addressed and resolved. The only failing check, JS Tests (shard 2/4) (work_queue_documentation.test.cjs), is unrelated to this PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

audit: sub-agent rows hide failures, ignore model aliases and dated ids, and have no per-agent spend (follow-up to #66740)

5 participants