Skip to content

💳 fix: Check Balance and Estimate Missing Usage for Remote Agent API Runs - #16657

Open
Odrec wants to merge 2 commits into
LibreChat-AI:devfrom
Odrec:fix/remote-agent-balance-check
Open

Odrec wants to merge 2 commits into
LibreChat-AI:devfrom
Odrec:fix/remote-agent-balance-check

Conversation

@Odrec

@Odrec Odrec commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The two remote agent endpoints, POST /api/agents/v1/chat/completions and POST /api/agents/v1/responses, differ from a chat turn in two ways that let API usage bypass balances.

First, they record a run's token usage afterwards but never admit the run against the user's balance first. The chat UI reserves the estimated prompt cost before every turn (BaseClient → checkBalance), and so do the Assistants endpoints, but the remote agent controllers never call it. With balance.enabled, a user whose credits are spent can keep calling agents with an API key, and the spend is clamped at 0 instead of being refused. A due auto-refill is applied inside that reservation, so API-only usage also never triggers a refill.

Second, they have no fallback when the provider reports no usage. createRun turns streamUsage off for custom endpoints unless it is set explicitly, so an OpenAI-compatible server such as vLLM behind LiteLLM streams without stream_options.include_usage and never returns token counts. The chat UI then falls back to a tokenizer estimate (getTokenCountForResponse → recordTokenUsage). The remote controllers return usage of 0 and write no transactions, so those runs are neither billed nor counted.

Both controllers now reserve the estimated prompt cost of the primary agent before the run starts. A user without enough credits gets an OpenAI-style 429 with type and code insufficient_quota before any response bytes are sent, including for streaming requests, so OpenAI SDKs report it as a quota error. The reservation is released once the run's usage has been recorded, or when the run fails. After the run, if no model call reported usage, a tokenizer estimate is added before usage is recorded, so the transactions and the response's usage reflect the run. Usage a provider reported is left as it is. With balances disabled, the admission step does nothing.

How it works

Both helpers are new, in packages/api/src/agents/remoteBalance.ts:

  • reserveRemoteAgentBalance is a thin wrapper around the existing checkBalance. It returns without reserving when balances are disabled. Otherwise it estimates the prompt as the token count of the agent's instructions plus the text of the request messages, and reserves that amount priced for the primary agent's model and endpoint.
  • addEstimatedUsageIfUnreported does nothing if any collected usage has input or output tokens. Otherwise it appends one usage entry: input is the same prompt estimate, output is the token count of the text, reasoning and tool-call arguments of the AI messages in run.getRunMessages(), which holds only the messages the run produced.
executeAgentRun
  execute
    initializeAgent, discovery, assertModelBoundContent
    reserveRemoteAgentBalance                         # token_balance → 429 insufficient_quota
    streaming headers, createRun, processStream
    addEstimatedUsageIfUnreported(collectedUsage)     # only when nothing was reported
    usageRecording = execution.track(recordCollectedUsage(...))
  beforeSettle
    execution.track(usageRecording → reservation.release())   # also runs when the run failed
  settle                                              # awaits tracked writes

In the OpenAI-compatible controller the balance check runs before the streaming headers are flushed. In the Responses controller it runs after the content checks, on the merged history (previous response messages plus the new input), which is also before streaming starts. The Responses controller adds the usage estimate on both its streaming and non-streaming paths.

Type of change

  • Bug fix

Testing

Tested environments/configuration: the helpers and controllers are covered by automated tests only, on macOS with Node 24. The zero-usage behavior was observed on a v0.8.8 deployment whose custom endpoints are vLLM behind LiteLLM: an API-key run returned usage of 0 and wrote no transactions, while a google agent on the same deployment reported usage and was recorded. Setting streamUsage: true on the custom endpoint (via addParams) made the same server return real counts, which is the provider-side workaround; this PR covers deployments that don't set it.

Automated tests:

  • packages/api/src/agents/remoteBalance.spec.ts:
    • The balance tests run against a real in-memory balance store. They check that a user whose credits are spent is refused, that the estimated prompt cost is held until the reservation is released, that a due auto-refill is applied before the request is admitted, and that nothing happens when balances are disabled.
    • The usage tests check that an estimate is added when nothing was reported, counting answers, reasoning and tool-call arguments but not tool results. Usage reported as all zeros counts as unreported, and reported usage is left untouched.
  • api/server/controllers/agents/__tests__/openai.spec.js and responses.unit.spec.js:
    • The reservation is made before createRun with the primary agent's model.
    • A refusal returns 429 insufficient_quota without starting the run or the stream, for streaming and non-streaming requests.
    • The reservation is released after usage is recorded, and when the run fails.
    • The usage estimate runs on the same collectedUsage array before recordCollectedUsage, in both controllers and on both Responses paths.
  • npx tsc --noEmit in packages/api and ESLint on the changed files pass.
  • packages/api src/agents suite: 4686 tests passed. The three SIGKILL escalation tests in src/agents/hooks/executor.spec.ts and src/agents/hooks/reaper.spec.ts fail the same way on unmodified dev on macOS. api server/controllers/agents and server/routes/agents: 1304 tests passed.

Screenshots / recordings

No user-facing change. The visible differences for API clients are the new 429 error response and non-zero usage from providers that report none.

Risk / compatibility

  • When balance.enabled is true, API requests that the balance cannot cover now fail with 429 instead of running. Deployments where API usage has effectively been unmetered will start seeing refusals.
  • Runs against providers that report no usage now write transactions and debit balances. They used to cost nothing.
  • Both estimates use the same tokenizer as the chat UI's fallback. The prompt estimate covers the primary agent's instructions and the message text, without tool definitions, attached file content or the extra prompts of a multi-step tool loop. So it is a lower bound, as the UI's is.
  • Found on a university deployment where every user can create API keys: the weekly per-user budget held in the chat UI but not for calls made with an API key.

🤖 Generated with Claude Code

The OpenAI-compatible (`/api/agents/v1/chat/completions`) and Open
Responses (`/api/agents/v1/responses`) endpoints recorded token usage
after a run but never admitted the run against the user's balance first.
A user whose credits were spent could keep calling agents with an API
key, and because the reservation is also where a due auto-refill is
applied, API-only usage never triggered a refill.

Both controllers now reserve the estimated prompt cost of the primary
agent with `reserveRemoteAgentBalance` before the run starts, as the
chat UI does for every turn. A user without enough credits gets an
OpenAI-style `429 insufficient_quota` error before any response bytes
are sent, including for streaming requests. The reservation is released
once the run's usage has been recorded, or when the run fails.
@codegraph-librechat codegraph-librechat Bot added the 🗺️ Backend Platform codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) label Oct 2, 2026
A provider that reports no token usage, such as an OpenAI-compatible
server streaming without `stream_options.include_usage`, left remote
agent runs with nothing to record: `/v1/chat/completions` and
`/v1/responses` returned zero usage and wrote no transactions, while the
chat UI falls back to a tokenizer estimate of the turn.

After the run, `addEstimatedUsageIfUnreported` adds an estimate when no
model call reported usage: the prompt is the agent instructions plus
the request messages, the output is the answers, reasoning and
tool-call arguments of the AI messages the run produced. Usage a
provider reported is left as it is.
@Odrec Odrec changed the title 💳 fix: Check Balance Before Remote Agent API Runs 💳 fix: Check Balance and Estimate Missing Usage for Remote Agent API Runs Oct 2, 2026
@codegraph-librechat codegraph-librechat Bot added 🗺️ Billing Accounts codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) 🗺️ Billing Engine codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) labels Oct 3, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

🗺️ Backend Platform codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) 🗺️ Billing Accounts codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) 🗺️ Billing Engine codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant