You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
audit: model_routing.classifier_cost is 0 and per-agent credit reconciliation warns on routed runs (regression from #66961) #67079
Since #66961 (fixes #66960), gh aw audit --json reports model_routing.classifier_cost as 0 requests and 0 credits on routed runs. On routed Copilot runs that have per-agent usage, it also emits a spurious per-agent AI credits (…) differ from non-classifier proxy total (…) warning.
The classifier request is still in the raw proxy log with "purpose": "routing_classification". Audit now takes per-request usage from usage/aw_session.jsonl, and the firewall.token_usage events in these runs' unified files carry no purpose, path or x_initiator.
Evidence
Three runs in githubnext/gh-aw-routing-sandbox, each audited into a fresh -o folder with c2f0000 (current main). Run 37859425051 was also audited with 5a8304e, the commit just before #66961.
per-agent AI credits (2.743) differ from non-classifier proxy total (2.635) for endpoint /responses
37859425051: selected (22.519) + deviated (0) + classifier (0) = 22.519, but the proxy total is 23.295. The missing 0.776 is the classifier request. At 5a8304e the three buckets sum exactly to the total.
37813733289: per-agent credits are main 15.227 + file-summarizer 1.791 = 17.018. The proxy total minus the 0.323 classifier request is also 17.018, so the residual is 0.000. The only cause of this warning is that the classifier is not subtracted.
37813720288: even with the classifier (0.060) excluded, the non-classifier total is 2.576, so this run still warns after the fix (residual 0.167). The agent-reported main credits (1.599) include 26,674 cache-write tokens that the proxy reports as 0, and quick-checker is 1.144 by the agent versus 1.110 by the proxy. That is a separate agent-versus-proxy pricing difference and is out of scope here. Do not use this run as the "no warning" acceptance case.
In all three usage/aw_session.jsonl files, no firewall.token_usage event has purpose or path. The raw sandbox/firewall/logs/api-proxy-logs/token-usage.jsonl downloaded with them has both, plus x_initiator.
Root cause
Model-routing costs.Report sub-agent failures, resolved models, and per-agent spend #66961 changed tokenUsageEntriesForRun in pkg/cli/model_routing.go. It used to read the raw proxy log through findTokenUsageFile and scanTokenUsageEntries; it now calls readUnifiedTokenUsageEntries in pkg/cli/token_usage_subagent_session.go. aggregateModelRoutingCosts counts a request as classifier traffic only when Purpose is routing_classification, so a unified event with no purpose lands in no bucket. The firewall_token_usage totals are still built from the raw log by parseTokenUsageFile in pkg/cli/token_usage_parse.go. Totals and buckets now come from different sources and no longer add up.
Credit reconciliation.reconcileAgentUsageCredits in pkg/cli/token_usage_subagent.go, called from augmentSubagentModelAttribution, subtracts classifier credits from TotalAIC only for unified entries whose purpose is routing_classification. With no purpose, the "non-classifier proxy total" still includes the classifier.
Endpoint in the warning. When no unified entry has a path, the warning falls back to summary.endpoint. buildTokenUsageSummary sets that from the first raw entry, which on routed runs is the classifier request. On the two Copilot runs this happens to equal the agent endpoint. On 37859425051, the classifier went to /chat/completions and every agent request went to /v1/messages?beta=true, so a warning on a routed Claude run would name the classifier's endpoint (/chat/completions or /responses) instead of /v1/messages. This was not reproduced as a warning on 37859425051 itself, because reconciliation only runs when per-agent usage exists.
Projection and spec.Report sub-agent failures, resolved models, and per-agent spend #66961 added purpose and path (with an endpoint alias) to the firewall.token_usage field map in actions/setup/js/unified_session_payload.cjs. Running the raw classifier record from 37859425051 through c2f0000's normalizeUnifiedSessionEvent keeps both fields. Three gaps remain:
The projection still drops x_initiator, and neither unifiedTokenUsageEntry nor TokenUsageEntry (pkg/cli/token_usage_types.go) has a field for it.
docs/src/content/docs/specs/unified-agent-session-specification.md lists firewall.token_usage only in the event-type table and never defines its payload fields. Nothing tells consumers they can rely on purpose, path or x_initiator.
Tests. The unified-path test for analyzeModelRouting in pkg/cli/model_routing_test.go has no classifier event. The reconcileAgentUsageCredits tests in pkg/cli/token_usage_declared_subagents_test.go build TokenUsageEntry values directly. No test exercises a real routed unified file end to end.
Proposed fix
This follows the maintainer guidance: everything audit and logs need must be in the unified agent session file, and consumers read it from there.
1. Unified projection and spec
In actions/setup/js/unified_session_payload.cjs, add xInitiator to the firewall.token_usage field map, which usage.report aliases. Accept the source names x_initiator and xInitiator, next to the existing purpose and path. Optionally, normalize path to the request pathname without the query string (the proxy records /v1/messages?beta=true) so endpoint comparisons are stable.
In the unified session spec, define the payload of firewall.token_usage and usage.report:
provider and model
purpose, for example agent, subagent or routing_classification
path, the request endpoint
xInitiator
requestId, status and durationMs
aic, totalAic and premiumRequests
usage
State that consumers identify router and classifier traffic by purpose, and that a missing purpose means unknown, not agent. Add a numbered T-UAS requirement for this. Update the "Audit and logs read sub-agent lifecycle and model evidence…" paragraph to say that classifier credits are identified by purpose and excluded from per-agent reconciliation.
Extend the projection tests in actions/setup/js so a raw record with purpose, path and x_initiator keeps all three fields.
2. Audit readers
Have unifiedTokenUsageEntry read xInitiator into a new TokenUsageEntry field. Give the field the JSON tag x_initiator so raw-log parsing fills it too.
Add a fallback for older unified files that lack purpose. Put it in readUnifiedTokenUsageEntries, or in a helper it calls, so that tokenUsageEntriesForRun and augmentSubagentModelAttribution both benefit. When the agent-phase firewall.token_usage events have no purpose, fill the missing purpose, path and x_initiator from one of these sources, in order:
The raw token-usage.jsonl, when it is in the run directory (located with findTokenUsageFile), joined on requestId. It was downloaded for all three runs above.
Otherwise, the unified file's own firewall.model_routing events. The classifier request is the token-usage event that meets all of these conditions:
its model matches the wire form of the classification record's classifier_model;
its timestamp falls between the stage: classification record and the next stage: selection record, allowing for its durationMs;
its requestId matches no stage: request record's request_id.
On 37859425051 this picks request 9dcb93e6-…: classification at 23:28:34.976, usage at 23:28:36.228, selection at 23:28:36.235. Repeat for each attempt when classifier_attempts is greater than 1.
The fallback only fills missing fields and never overrides a value present in the unified event. It logs which source supplied the value through tokenUsageSubagentLog.
In reconcileAgentUsageCredits, take endpoint names only from non-classifier entries. When no entry has a path, do not fall back blindly to summary.endpoint. Either make buildTokenUsageSummary skip routing_classification entries when it picks summary.endpoint, or report unknown endpoint.
aggregateModelRoutingCosts needs no change once entries carry purpose, because it already buckets on that field.
3. Regression tests
Routed fixture. Add a redacted fixture under pkg/cli/testdata/ modelled on 37813733289. It contains:
a usage/aw_session.jsonl with a routing_classificationfirewall.token_usage event, the main and sub-agent requests, and matching per-agent usage (session.shutdown and subagent.*);
a usage/agent/model-routing.jsonl with classification, selection and request records.
Through analyzeModelRouting, assert that classifier_cost is 1 request with the classifier's credits, and that classifier + selected + deviated equals the proxy total. Through augmentSubagentModelAttribution, assert that no reconciliation warning is emitted.
Legacy variant. Use the same fixture, but remove purpose and path from the unified events, as in the three runs above. Assert the same results twice: once through the raw-log fallback, and once with the raw log removed, through the firewall.model_routing timing join.
Claude variant. Modelled on 37859425051: the classifier is on /chat/completions and the agent is on /v1/messages. Assert that any reconciliation warning names the agent endpoint, never the classifier's.
JS projection test. As in section 1.
Acceptance
gh aw audit 37859425051 --json reports model_routing.classifier_cost as 1 request and 0.776 credits, matching 5a8304e. Classifier + selected + deviated equals 23.295.
gh aw audit 37813733289 --json reports classifier_cost as 1 request and 0.323 credits, with no per-agent AI credits … differ warning.
gh aw audit 37813720288 --json reports classifier_cost as 1 request and 0.060 credits. Any remaining reconciliation warning on this run comes from the separate agent-versus-proxy pricing residual (0.167) described above.
Summary
Since #66961 (fixes #66960),
gh aw audit --jsonreportsmodel_routing.classifier_costas 0 requests and 0 credits on routed runs. On routed Copilot runs that have per-agent usage, it also emits a spuriousper-agent AI credits (…) differ from non-classifier proxy total (…)warning.The classifier request is still in the raw proxy log with
"purpose": "routing_classification". Audit now takes per-request usage fromusage/aw_session.jsonl, and thefirewall.token_usageevents in these runs' unified files carry nopurpose,pathorx_initiator.Evidence
Three runs in githubnext/gh-aw-routing-sandbox, each audited into a fresh
-ofolder with c2f0000 (current main). Run 37859425051 was also audited with 5a8304e, the commit just before #66961.total_aictoken-usage.jsonlclassifier_cost@ 5a8304eclassifier_cost@ c2f0000claude-opus-5,/chat/completions)claude-sonnet-5,/chat/completions)per-agent AI credits (17.018) differ from non-classifier proxy total (17.341) for endpoint /chat/completionsgpt-5.6-luna,/responses)per-agent AI credits (2.743) differ from non-classifier proxy total (2.635) for endpoint /responsesusage/aw_session.jsonlfiles, nofirewall.token_usageevent haspurposeorpath. The rawsandbox/firewall/logs/api-proxy-logs/token-usage.jsonldownloaded with them has both, plusx_initiator.Root cause
tokenUsageEntriesForRuninpkg/cli/model_routing.go. It used to read the raw proxy log throughfindTokenUsageFileandscanTokenUsageEntries; it now callsreadUnifiedTokenUsageEntriesinpkg/cli/token_usage_subagent_session.go.aggregateModelRoutingCostscounts a request as classifier traffic only whenPurposeisrouting_classification, so a unified event with nopurposelands in no bucket. Thefirewall_token_usagetotals are still built from the raw log byparseTokenUsageFileinpkg/cli/token_usage_parse.go. Totals and buckets now come from different sources and no longer add up.reconcileAgentUsageCreditsinpkg/cli/token_usage_subagent.go, called fromaugmentSubagentModelAttribution, subtracts classifier credits fromTotalAIConly for unified entries whosepurposeisrouting_classification. With nopurpose, the "non-classifier proxy total" still includes the classifier.path, the warning falls back tosummary.endpoint.buildTokenUsageSummarysets that from the first raw entry, which on routed runs is the classifier request. On the two Copilot runs this happens to equal the agent endpoint. On 37859425051, the classifier went to/chat/completionsand every agent request went to/v1/messages?beta=true, so a warning on a routed Claude run would name the classifier's endpoint (/chat/completionsor/responses) instead of/v1/messages. This was not reproduced as a warning on 37859425051 itself, because reconciliation only runs when per-agent usage exists.purposeandpath(with anendpointalias) to thefirewall.token_usagefield map inactions/setup/js/unified_session_payload.cjs. Running the raw classifier record from 37859425051 through c2f0000'snormalizeUnifiedSessionEventkeeps both fields. Three gaps remain:x_initiator, and neitherunifiedTokenUsageEntrynorTokenUsageEntry(pkg/cli/token_usage_types.go) has a field for it.docs/src/content/docs/specs/unified-agent-session-specification.mdlistsfirewall.token_usageonly in the event-type table and never defines its payload fields. Nothing tells consumers they can rely onpurpose,pathorx_initiator.analyzeModelRoutinginpkg/cli/model_routing_test.gohas no classifier event. ThereconcileAgentUsageCreditstests inpkg/cli/token_usage_declared_subagents_test.gobuildTokenUsageEntryvalues directly. No test exercises a real routed unified file end to end.Proposed fix
This follows the maintainer guidance: everything audit and logs need must be in the unified agent session file, and consumers read it from there.
1. Unified projection and spec
In
actions/setup/js/unified_session_payload.cjs, addxInitiatorto thefirewall.token_usagefield map, whichusage.reportaliases. Accept the source namesx_initiatorandxInitiator, next to the existingpurposeandpath. Optionally, normalizepathto the request pathname without the query string (the proxy records/v1/messages?beta=true) so endpoint comparisons are stable.In the unified session spec, define the payload of
firewall.token_usageandusage.report:providerandmodelpurpose, for exampleagent,subagentorrouting_classificationpath, the request endpointxInitiatorrequestId,statusanddurationMsaic,totalAicandpremiumRequestsusageState that consumers identify router and classifier traffic by
purpose, and that a missingpurposemeans unknown, not agent. Add a numbered T-UAS requirement for this. Update the "Audit and logs read sub-agent lifecycle and model evidence…" paragraph to say that classifier credits are identified bypurposeand excluded from per-agent reconciliation.Extend the projection tests in
actions/setup/jsso a raw record withpurpose,pathandx_initiatorkeeps all three fields.2. Audit readers
Have
unifiedTokenUsageEntryreadxInitiatorinto a newTokenUsageEntryfield. Give the field the JSON tagx_initiatorso raw-log parsing fills it too.Add a fallback for older unified files that lack
purpose. Put it inreadUnifiedTokenUsageEntries, or in a helper it calls, so thattokenUsageEntriesForRunandaugmentSubagentModelAttributionboth benefit. When the agent-phasefirewall.token_usageevents have nopurpose, fill the missingpurpose,pathandx_initiatorfrom one of these sources, in order:The raw
token-usage.jsonl, when it is in the run directory (located withfindTokenUsageFile), joined onrequestId. It was downloaded for all three runs above.Otherwise, the unified file's own
firewall.model_routingevents. The classifier request is the token-usage event that meets all of these conditions:classifier_model;stage: classificationrecord and the nextstage: selectionrecord, allowing for itsdurationMs;requestIdmatches nostage: requestrecord'srequest_id.On 37859425051 this picks request
9dcb93e6-…: classification at 23:28:34.976, usage at 23:28:36.228, selection at 23:28:36.235. Repeat for eachattemptwhenclassifier_attemptsis greater than 1.The fallback only fills missing fields and never overrides a value present in the unified event. It logs which source supplied the value through
tokenUsageSubagentLog.In
reconcileAgentUsageCredits, take endpoint names only from non-classifier entries. When no entry has a path, do not fall back blindly tosummary.endpoint. Either makebuildTokenUsageSummaryskiprouting_classificationentries when it pickssummary.endpoint, or reportunknown endpoint.aggregateModelRoutingCostsneeds no change once entries carrypurpose, because it already buckets on that field.3. Regression tests
Routed fixture. Add a redacted fixture under
pkg/cli/testdata/modelled on 37813733289. It contains:usage/aw_session.jsonlwith arouting_classificationfirewall.token_usageevent, the main and sub-agent requests, and matching per-agent usage (session.shutdownandsubagent.*);usage/agent/model-routing.jsonlwith classification, selection and request records.Through
analyzeModelRouting, assert thatclassifier_costis 1 request with the classifier's credits, and that classifier + selected + deviated equals the proxy total. ThroughaugmentSubagentModelAttribution, assert that no reconciliation warning is emitted.Legacy variant. Use the same fixture, but remove
purposeandpathfrom the unified events, as in the three runs above. Assert the same results twice: once through the raw-log fallback, and once with the raw log removed, through thefirewall.model_routingtiming join.Claude variant. Modelled on 37859425051: the classifier is on
/chat/completionsand the agent is on/v1/messages. Assert that any reconciliation warning names the agent endpoint, never the classifier's.JS projection test. As in section 1.
Acceptance
gh aw audit 37859425051 --jsonreportsmodel_routing.classifier_costas 1 request and 0.776 credits, matching 5a8304e. Classifier + selected + deviated equals 23.295.gh aw audit 37813733289 --jsonreportsclassifier_costas 1 request and 0.323 credits, with noper-agent AI credits … differwarning.gh aw audit 37813720288 --jsonreportsclassifier_costas 1 request and 0.060 credits. Any remaining reconciliation warning on this run comes from the separate agent-versus-proxy pricing residual (0.167) described above.