Conversation
The OpenAI-compatible (`/api/agents/v1/chat/completions`) and Open Responses (`/api/agents/v1/responses`) endpoints recorded token usage after a run but never admitted the run against the user's balance first. A user whose credits were spent could keep calling agents with an API key, and because the reservation is also where a due auto-refill is applied, API-only usage never triggered a refill. Both controllers now reserve the estimated prompt cost of the primary agent with `reserveRemoteAgentBalance` before the run starts, as the chat UI does for every turn. A user without enough credits gets an OpenAI-style `429 insufficient_quota` error before any response bytes are sent, including for streaming requests. The reservation is released once the run's usage has been recorded, or when the run fails.
A provider that reports no token usage, such as an OpenAI-compatible server streaming without `stream_options.include_usage`, left remote agent runs with nothing to record: `/v1/chat/completions` and `/v1/responses` returned zero usage and wrote no transactions, while the chat UI falls back to a tokenizer estimate of the turn. After the run, `addEstimatedUsageIfUnreported` adds an estimate when no model call reported usage: the prompt is the agent instructions plus the request messages, the output is the answers, reasoning and tool-call arguments of the AI messages the run produced. Usage a provider reported is left as it is.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The two remote agent endpoints,
POST /api/agents/v1/chat/completionsandPOST /api/agents/v1/responses, differ from a chat turn in two ways that let API usage bypass balances.First, they record a run's token usage afterwards but never admit the run against the user's balance first. The chat UI reserves the estimated prompt cost before every turn (
BaseClient→checkBalance), and so do the Assistants endpoints, but the remote agent controllers never call it. Withbalance.enabled, a user whose credits are spent can keep calling agents with an API key, and the spend is clamped at 0 instead of being refused. A due auto-refill is applied inside that reservation, so API-only usage also never triggers a refill.Second, they have no fallback when the provider reports no usage.
createRunturnsstreamUsageoff for custom endpoints unless it is set explicitly, so an OpenAI-compatible server such as vLLM behind LiteLLM streams withoutstream_options.include_usageand never returns token counts. The chat UI then falls back to a tokenizer estimate (getTokenCountForResponse→recordTokenUsage). The remote controllers returnusageof 0 and write no transactions, so those runs are neither billed nor counted.Both controllers now reserve the estimated prompt cost of the primary agent before the run starts. A user without enough credits gets an OpenAI-style
429with type and codeinsufficient_quotabefore any response bytes are sent, including for streaming requests, so OpenAI SDKs report it as a quota error. The reservation is released once the run's usage has been recorded, or when the run fails. After the run, if no model call reported usage, a tokenizer estimate is added before usage is recorded, so the transactions and the response'susagereflect the run. Usage a provider reported is left as it is. With balances disabled, the admission step does nothing.How it works
Both helpers are new, in
packages/api/src/agents/remoteBalance.ts:reserveRemoteAgentBalanceis a thin wrapper around the existingcheckBalance. It returns without reserving when balances are disabled. Otherwise it estimates the prompt as the token count of the agent's instructions plus the text of the request messages, and reserves that amount priced for the primary agent's model and endpoint.addEstimatedUsageIfUnreporteddoes nothing if any collected usage has input or output tokens. Otherwise it appends one usage entry: input is the same prompt estimate, output is the token count of the text, reasoning and tool-call arguments of the AI messages inrun.getRunMessages(), which holds only the messages the run produced.In the OpenAI-compatible controller the balance check runs before the streaming headers are flushed. In the Responses controller it runs after the content checks, on the merged history (previous response messages plus the new input), which is also before streaming starts. The Responses controller adds the usage estimate on both its streaming and non-streaming paths.
Type of change
Testing
Tested environments/configuration: the helpers and controllers are covered by automated tests only, on macOS with Node 24. The zero-usage behavior was observed on a v0.8.8 deployment whose custom endpoints are vLLM behind LiteLLM: an API-key run returned
usageof 0 and wrote no transactions, while agoogleagent on the same deployment reported usage and was recorded. SettingstreamUsage: trueon the custom endpoint (viaaddParams) made the same server return real counts, which is the provider-side workaround; this PR covers deployments that don't set it.Automated tests:
packages/api/src/agents/remoteBalance.spec.ts:api/server/controllers/agents/__tests__/openai.spec.jsandresponses.unit.spec.js:createRunwith the primary agent's model.429 insufficient_quotawithout starting the run or the stream, for streaming and non-streaming requests.collectedUsagearray beforerecordCollectedUsage, in both controllers and on both Responses paths.npx tsc --noEmitinpackages/apiand ESLint on the changed files pass.packages/apisrc/agentssuite: 4686 tests passed. The threeSIGKILLescalation tests insrc/agents/hooks/executor.spec.tsandsrc/agents/hooks/reaper.spec.tsfail the same way on unmodifieddevon macOS.apiserver/controllers/agentsandserver/routes/agents: 1304 tests passed.Screenshots / recordings
No user-facing change. The visible differences for API clients are the new
429error response and non-zerousagefrom providers that report none.Risk / compatibility
balance.enabledis true, API requests that the balance cannot cover now fail with429instead of running. Deployments where API usage has effectively been unmetered will start seeing refusals.🤖 Generated with Claude Code