refactor(attribution)!: remove project_id and the LiteLLM cost attribution built on it - #45
polakamtejas wants to merge 3 commits into
Conversation
…ution built on it project_id was the hosted platform's spend key: `agent-env task create` required it, and it became the LiteLLM `user`, a `projectId:` tag, the Modal App name, and a field the A2A validator checked agents forward. agent-env has no project concept, so it is removed rather than generalized. Sandbox cost dimensions stay in the open `attribution` dict that providers already read. - CLI: `task create` no longer takes `--project-id`; `task run` and `run-batch` drop the option and the missing-attribution banner. - Task: no `project_id` field. A stored document that still carries the key loads, and the key is not written back. - LiteLLM: `utils/litellm_attribution.py` is gone. The rubrics judge's direct call sends no `user` or tags, and the auto-deployed judge gets no LITELLM_PROJECT_ID / LITELLM_USER (customer_id) env vars or attribution. - Modal / E2B: one Modal App, the configured base name, instead of one per project; `project_id` is no longer an App tag or E2B metadata key. - Agents: prompt_agent and the judge stop sending `project_id` through agent-config (`task_id` is still sent). The A2A validator drops the static and runtime LiteLLM-attribution checks, and seeds `task_id` with the validation task's own id. - Explorer: `RunRequest.project_id` is gone, so a request that sends it now gets a 422 (the model forbids extra keys); the UI stops sending it. - Tests and the protocol README use neutral attribution keys. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| # Modal billing breaks cost down by App tags; sandbox-level tags don't show up in billing. | ||
| app_tags = _build_cost_attribution_tags(attribution) | ||
| app_name = _app_name_for_project(self._app_name, app_tags.get("project_id")) | ||
| app_name = self._app_name |
There was a problem hiding this comment.
When runs have different customer or team values, they now share one Modal App. _get_app sets that App’s tags only on the first lookup, so later sandboxes can have their costs assigned to the first run’s customer or team. The Modal VM provider has the same issue. Keep separate billing buckets for different attribution values.
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/providers/sandbox_providers/modal_sandbox.py
Line: 258
Comment:
**Costs get the wrong tags**
When runs have different `customer` or `team` values, they now share one Modal App. `_get_app` sets that App’s tags only on the first lookup, so later sandboxes can have their costs assigned to the first run’s customer or team. The Modal VM provider has the same issue. Keep separate billing buckets for different attribution values.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.There was a problem hiding this comment.
we are removing cusotmer/team vsalues
There was a problem hiding this comment.
If customer and team attribution are also being removed, then both findings are no longer applicable: there would be no per-run customer/team values to misattribute, and the judge would not need to preserve customer_id.
However, in the revision reviewed, _build_cost_attribution_tags() still emits customer and team, and the PR description explicitly says those tags remain. Therefore, the findings are valid for the current revision, but should be dismissed if a follow-up change removes those dimensions everywhere.
Where project_id was a sample key beside others, the key is removed rather than swapped for cost_center: the attribution-probe and caller-metadata checks, the e2b attribution test and the chain-forwarding test. The test that cost_center is not a Modal tag, and the assertion that agent-config rejects the key, only existed for project_id and are removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
🤖 Automated Claude routine — Env Pod PR Context Review. This is a bot, not Edgar Arakelyan; not a human review. Offering one piece of context that an in-repo
Because Worth either a line in the break list or a tolerated-and-ignored Generated by Claude Code |
…B metadata Replace the hardcoded product/customer/team App tags and the pipeline_step/run_id-only sandbox tags with one generic _attribution_tags. The shared Modal App is tagged once, so it carries only the deployment's [sandbox.attribution]; each sandbox carries the run's full attribution filled from those defaults. E2B metadata is the same resolved dict. A key Modal can't take raises instead of being rewritten, and the sandbox-started log nests the tags so a key like `name` can't clobber a LogRecord field. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| metadata = { | ||
| key: value | ||
| for key, value in apply_default_attribution(dict(attribution or {})).items() | ||
| if value is not None | ||
| } |
There was a problem hiding this comment.
Unexposed ports return on reconnect
If a caller sets agent_env_exposed_ports as an attribution key and creates an E2B sandbox with no exposed ports, this code keeps the caller’s value in the sandbox metadata. On reconnect, the provider reads it as its saved port list and adds those ports to tunnel_urls. Code that relies on that list can then treat ports the caller never exposed as available. Keep this provider-owned key out of attribution metadata.
| metadata = { | |
| key: value | |
| for key, value in apply_default_attribution(dict(attribution or {})).items() | |
| if value is not None | |
| } | |
| metadata = { | |
| key: value | |
| for key, value in apply_default_attribution(dict(attribution or {})).items() | |
| if value is not None | |
| } | |
| metadata.pop(_EXPOSED_PORTS_METADATA_KEY, None) |
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/providers/sandbox_providers/e2b/provider.py
Line: 146-150
Comment:
**Unexposed ports return on reconnect**
If a caller sets `agent_env_exposed_ports` as an attribution key and creates an E2B sandbox with no exposed ports, this code keeps the caller’s value in the sandbox metadata. On reconnect, the provider reads it as its saved port list and adds those ports to `tunnel_urls`. Code that relies on that list can then treat ports the caller never exposed as available. Keep this provider-owned key out of attribution metadata.
```suggestion
metadata = {
key: value
for key, value in apply_default_attribution(dict(attribution or {})).items()
if value is not None
}
metadata.pop(_EXPOSED_PORTS_METADATA_KEY, None)
```
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
project_idwas the hosted platform's cost-attribution key.agent-env task createrequired it, and it then turned into:userfield and aprojectId:tag on judge callsagent-env has no project concept, so this removes it rather than generalizing it. Sandbox cost dimensions keep flowing through the open
attributiondict (agent_env/attribution.py). That path is now generic too: no key is hardcoded, and providers carry every key.What's removed
agent-env task createno longer takes--project-id(it was required).task runandrun-batchdrop the option and the yellow "LiteLLM cost attribution is incomplete" banner, plus the step-type list that only fed it.Taskhas noproject_idfield. A stored task document that still carries the key loads fine, and the key is not written back (newtst/unit/task/task_document_test.py).utils/litellm_attribution.pyis deleted. Its one caller, the rubrics judge's direct LiteLLM call, now sends nouserormetadata.tags. The helper only produced output when aproject_idwas set, so withproject_idgone it had nothing left to do, which is why I removed it instead of keeping ataskId-only tag. The auto-deployed judge agent no longer getsLITELLM_PROJECT_IDorLITELLM_USER.LITELLM_USERcame fromcontext.metadata["customer_id"], the same platform attribution, which nothing in this repo sets. The judge also no longer gets aproject_id/customerattribution dict.app_name/AGENT_ENV_MODAL_APP_NAME), instead of<base>-<project_id>. Attribution no longer hardcodesproduct/customer/team; one_attribution_tagsreplaces both tag builders:[sandbox.attribution]defaults, under any keys. The App is shared and tagged once, so a run's own values can't go there without billing later runs to the first run's tags.attributiondict (any keys,pipeline_step/run_idincluded), with unset keys filled from those defaults.[a-zA-Z0-9._-]) raises aValueErrorbefore the App is looked up, instead of being rewritten. Values are sanitized as before.modal_sandbox_startedlog line nests the tags undermodal_sandbox_tagsand keepspipeline_step/run_idtop-level for the container join, so a key likenamecan't clobber a LogRecord field.prompt_agentand the rubrics judge stop sendingproject_idthrough/ext/agent-config.task_idis still sent. The test echo agent drops the config field.verify_a2a_litellm_attributionand runtimeverify_a2a_litellm_attribution_runtimechecks are removed. Both checked LiteLLM spend attribution: the runtime one planted a probeproject_idandtask_idand read them back over/ext/attribution-probe. What's left of them, whether an agent passestask_idon to its model gateway, only means something for spend attribution. Agent-config itself stays covered by the identity check, which sets config through the same extension and checks that the served card changes. The validator now seedstask_idwith the validation task's own id.agent-env a2a-agent validatestops printing the twovalidated_litellm_attribution*blocks. The protocol package'surn:agentenv:attribution-probe/v1extension stays: it is an open, key-agnostic map, and only its README example and test values namedproject_id.RunRequest.project_idis removed, and the UI stops sending it: the runner's?projectId=URL parameter, and the Start Runs panel's copy oftask.project_id.Breaking for downstream
--project-idontask create,task runandrun-batchnow fails with "No such option". Scripts that pass it must drop it.POST /api/v1/tasks/{id}/runand/runswithproject_idin the body now return 422, becauseRunRequestforbids extra keys.user/projectId:/taskId:tags. A proxy that attributes or gates spend on them no longer gets them from agent-env.[sandbox.attribution]: a run's ownproduct/customer/teammove to the sandbox tags, which Modal's billing report doesn't break down by. An attribution key that isn't a valid Modal tag key now fails the create.project_idand carries every attribution key, includingpipeline_step/run_id.project_idthrough agent-config.validated_litellm_attribution*agent metadata keys are gone.check_plugin_api.pyreports no break to the plugin surface. The!marks the CLI and API breaks above.Testing
pytest tst/unit packages/agentenv-protocol/tests -n auto: 5783 passed. The only failures are the 2test_mcp_cli_buildertests that fail identically on main here, because the generated CLI's interpreter has noclick.attribution_wire_test.py), and a tag namednamedoesn't clobber the log record (modal_sandbox_started_log_test.py).npx tsc --noEmit,npm run test:smoke(13 ran, the 2 skips need trajectory fixtures) andnpm run build:staticpass.agent-env task create t.json --id live-checkwith no flag succeeds.--project-id xis rejected ("No such option").agent-env task run --id live-checkscores 1.0, and its context has noproject_id.agent-env run hellopasses.git grep project_idnow finds only GCP's unrelatedquota_project_id(gcs_object_store.pyand its test) and seed data for the mock Jira and Figma services intst/integration/env/gateway/gateway_test.py.🤖 Generated with Claude Code
The PR does not appear safe to merge while Modal billing loses run-specific buckets and E2B can restore ports the caller never exposed.
Fix with agent prompt
Summary
AgentEnv removes
project_idas a built-in task and LiteLLM attribution field because the framework has no project concept. Sandbox providers still carry open attribution dimensions, and A2A validation no longer checks LiteLLM spend attribution.Reviews (3) · Last reviewed commit: "refactor(attribution): carry any attribu..."