Skip to content

refactor(attribution)!: remove project_id and the LiteLLM cost attribution built on it - #45

Open
polakamtejas wants to merge 3 commits into
mainfrom
tejas/remove-project-id
Open

polakamtejas wants to merge 3 commits into
mainfrom
tejas/remove-project-id

Conversation

@polakamtejas

@polakamtejas polakamtejas commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

project_id was the hosted platform's cost-attribution key. agent-env task create required it, and it then turned into:

  • the LiteLLM user field and a projectId: tag on judge calls
  • the Modal App name (one App per project)
  • an agent-config field sent to agents
  • two A2A validator checks that agents forward it

agent-env has no project concept, so this removes it rather than generalizing it. Sandbox cost dimensions keep flowing through the open attribution dict (agent_env/attribution.py). That path is now generic too: no key is hardcoded, and providers carry every key.

What's removed

  • CLI: agent-env task create no longer takes --project-id (it was required). task run and run-batch drop the option and the yellow "LiteLLM cost attribution is incomplete" banner, plus the step-type list that only fed it.
  • Task: Task has no project_id field. A stored task document that still carries the key loads fine, and the key is not written back (new tst/unit/task/task_document_test.py).
  • LiteLLM: utils/litellm_attribution.py is deleted. Its one caller, the rubrics judge's direct LiteLLM call, now sends no user or metadata.tags. The helper only produced output when a project_id was set, so with project_id gone it had nothing left to do, which is why I removed it instead of keeping a taskId-only tag. The auto-deployed judge agent no longer gets LITELLM_PROJECT_ID or LITELLM_USER. LITELLM_USER came from context.metadata["customer_id"], the same platform attribution, which nothing in this repo sets. The judge also no longer gets a project_id/customer attribution dict.
  • Modal / E2B: every sandbox goes under one Modal App, the configured base name (app_name / AGENT_ENV_MODAL_APP_NAME), instead of <base>-<project_id>. Attribution no longer hardcodes product / customer / team; one _attribution_tags replaces both tag builders:
    • App tags are only the deployment's [sandbox.attribution] defaults, under any keys. The App is shared and tagged once, so a run's own values can't go there without billing later runs to the first run's tags.
    • Sandbox tags are the run's whole attribution dict (any keys, pipeline_step / run_id included), with unset keys filled from those defaults.
    • E2B sandbox metadata is that same resolved dict, so E2B no longer imports a Modal helper.
    • A key Modal can't take as a tag (anything but 1–63 of [a-zA-Z0-9._-]) raises a ValueError before the App is looked up, instead of being rewritten. Values are sanitized as before.
    • The modal_sandbox_started log line nests the tags under modal_sandbox_tags and keeps pipeline_step / run_id top-level for the container join, so a key like name can't clobber a LogRecord field.
  • Agents: prompt_agent and the rubrics judge stop sending project_id through /ext/agent-config. task_id is still sent. The test echo agent drops the config field.
  • A2A validation: the static verify_a2a_litellm_attribution and runtime verify_a2a_litellm_attribution_runtime checks are removed. Both checked LiteLLM spend attribution: the runtime one planted a probe project_id and task_id and read them back over /ext/attribution-probe. What's left of them, whether an agent passes task_id on to its model gateway, only means something for spend attribution. Agent-config itself stays covered by the identity check, which sets config through the same extension and checks that the served card changes. The validator now seeds task_id with the validation task's own id. agent-env a2a-agent validate stops printing the two validated_litellm_attribution* blocks. The protocol package's urn:agentenv:attribution-probe/v1 extension stays: it is an open, key-agnostic map, and only its README example and test values named project_id.
  • Explorer: RunRequest.project_id is removed, and the UI stops sending it: the runner's ?projectId= URL parameter, and the Start Runs panel's copy of task.project_id.

Breaking for downstream

  • CLI: --project-id on task create, task run and run-batch now fails with "No such option". Scripts that pass it must drop it.
  • Explorer API: POST /api/v1/tasks/{id}/run and /runs with project_id in the body now return 422, because RunRequest forbids extra keys.
  • LiteLLM: requests from the rubrics judge carry no user / projectId: / taskId: tags. A proxy that attributes or gates spend on them no longer gets them from agent-env.
  • Modal: sandboxes land in one App (the base name) instead of per-project Apps. App tags are now only [sandbox.attribution]: a run's own product / customer / team move to the sandbox tags, which Modal's billing report doesn't break down by. An attribution key that isn't a valid Modal tag key now fails the create.
  • E2B: metadata loses project_id and carries every attribution key, including pipeline_step / run_id.
  • Agents: agents no longer receive project_id through agent-config.
  • A2A validation: the two LiteLLM attribution checks and their validated_litellm_attribution* agent metadata keys are gone.

check_plugin_api.py reports no break to the plugin surface. The ! marks the CLI and API breaks above.

Testing

  • pytest tst/unit packages/agentenv-protocol/tests -n auto: 5783 passed. The only failures are the 2 test_mcp_cli_builder tests that fail identically on main here, because the generated CLI's interpreter has no click.
  • Generic attribution tests: any key reaches the sandbox tags while the App gets only the defaults, a bad key fails before the App is looked up (attribution_wire_test.py), and a tag named name doesn't clobber the log record (modal_sandbox_started_log_test.py).
  • UI: npx tsc --noEmit, npm run test:smoke (13 ran, the 2 skips need trajectory fixtures) and npm run build:static pass.
  • Live check in a fresh local state:
    • agent-env task create t.json --id live-check with no flag succeeds.
    • --project-id x is rejected ("No such option").
    • agent-env task run --id live-check scores 1.0, and its context has no project_id.
    • agent-env run hello passes.
  • git grep project_id now finds only GCP's unrelated quota_project_id (gcs_object_store.py and its test) and seed data for the mock Jira and Figma services in tst/integration/env/gateway/gateway_test.py.

🤖 Generated with Claude Code

RetriggerConfidence Score: 3/5

The PR does not appear safe to merge while Modal billing loses run-specific buckets and E2B can restore ports the caller never exposed.

Fix All in CursorFindings

  1. P1 Unexposed ports return on reconnect ▶
  2. P1 Costs get the wrong tags ▶
Fix with agent prompt
### Issue 1
src/agent_env/providers/sandbox_providers/e2b/provider.py:146-150
If a caller sets `agent_env_exposed_ports` as an attribution key and creates an E2B sandbox with no exposed ports, this code keeps the caller’s value in the sandbox metadata. On reconnect, the provider reads it as its saved port list and adds those ports to `tunnel_urls`. Code that relies on that list can then treat ports the caller never exposed as available. Keep this provider-owned key out of attribution metadata.

```suggestion
        metadata = {
            key: value
            for key, value in apply_default_attribution(dict(attribution or {})).items()
            if value is not None
        }
        metadata.pop(_EXPOSED_PORTS_METADATA_KEY, None)
```

### Issue 2
src/agent_env/providers/sandbox_providers/modal_sandbox.py:undefined-261
When runs have different `customer` or `team` values, they now share one Modal App. `_get_app` sets that App’s tags only on the first lookup, so later sandboxes can have their costs assigned to the first run’s customer or team. The Modal VM provider has the same issue. Keep separate billing buckets for different attribution values.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

AgentEnv removes project_id as a built-in task and LiteLLM attribution field because the framework has no project concept. Sandbox providers still carry open attribution dimensions, and A2A validation no longer checks LiteLLM spend attribution.

  • Removes the project ID option and field from task creation, runs, Explorer requests, and agent configuration.
  • Sends sandbox attribution dimensions to Modal tags and E2B metadata, using one configured Modal App name.
  • Removes the LiteLLM attribution helper and the A2A checks built around it.

Reviews (3) · Last reviewed commit: "refactor(attribution): carry any attribu..."

…ution built on it

project_id was the hosted platform's spend key: `agent-env task create`
required it, and it became the LiteLLM `user`, a `projectId:` tag, the Modal
App name, and a field the A2A validator checked agents forward. agent-env has
no project concept, so it is removed rather than generalized. Sandbox cost
dimensions stay in the open `attribution` dict that providers already read.

- CLI: `task create` no longer takes `--project-id`; `task run` and
  `run-batch` drop the option and the missing-attribution banner.
- Task: no `project_id` field. A stored document that still carries the key
  loads, and the key is not written back.
- LiteLLM: `utils/litellm_attribution.py` is gone. The rubrics judge's
  direct call sends no `user` or tags, and the auto-deployed judge gets no
  LITELLM_PROJECT_ID / LITELLM_USER (customer_id) env vars or attribution.
- Modal / E2B: one Modal App, the configured base name, instead of one per
  project; `project_id` is no longer an App tag or E2B metadata key.
- Agents: prompt_agent and the judge stop sending `project_id` through
  agent-config (`task_id` is still sent). The A2A validator drops the static
  and runtime LiteLLM-attribution checks, and seeds `task_id` with the
  validation task's own id.
- Explorer: `RunRequest.project_id` is gone, so a request that sends it now
  gets a 422 (the model forbids extra keys); the UI stops sending it.
- Tests and the protocol README use neutral attribution keys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@polakamtejas
polakamtejas requested a review from a team as a code owner October 1, 2026 23:17
# Modal billing breaks cost down by App tags; sandbox-level tags don't show up in billing.
app_tags = _build_cost_attribution_tags(attribution)
app_name = _app_name_for_project(self._app_name, app_tags.get("project_id"))
app_name = self._app_name

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Costs get the wrong tags

When runs have different customer or team values, they now share one Modal App. _get_app sets that App’s tags only on the first lookup, so later sandboxes can have their costs assigned to the first run’s customer or team. The Modal VM provider has the same issue. Keep separate billing buckets for different attribution values.

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/providers/sandbox_providers/modal_sandbox.py
Line: 258

Comment:
**Costs get the wrong tags**

When runs have different `customer` or `team` values, they now share one Modal App. `_get_app` sets that App’s tags only on the first lookup, so later sandboxes can have their costs assigned to the first run’s customer or team. The Modal VM provider has the same issue. Keep separate billing buckets for different attribution values.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we are removing cusotmer/team vsalues

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If customer and team attribution are also being removed, then both findings are no longer applicable: there would be no per-run customer/team values to misattribute, and the judge would not need to preserve customer_id.

However, in the revision reviewed, _build_cost_attribution_tags() still emits customer and team, and the PR description explicitly says those tags remain. Therefore, the findings are valid for the current revision, but should be dismissed if a follow-up change removes those dimensions everywhere.

Comment thread src/agent_env/task_step/task_steps/verifiers/rubrics_verifier.py
Comment thread src/agent_env/a2a_agent/validator.py
Where project_id was a sample key beside others, the key is removed rather than
swapped for cost_center: the attribution-probe and caller-metadata checks, the
e2b attribution test and the chain-forwarding test. The test that cost_center
is not a Modal tag, and the assertion that agent-config rejects the key, only
existed for project_id and are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@earakely-scale

Copy link
Copy Markdown
Collaborator

🤖 Automated Claude routine — Env Pod PR Context Review. This is a bot, not Edgar Arakelyan; not a human review.

Offering one piece of context that an in-repo git grep can't surface: the "Breaking for downstream" list looks complete for the CLI, Explorer API, LiteLLM, Modal and A2A surfaces, but not for the Python Task API — which is the surface an embedding application uses.

  • src/agent_env/task/task.py:303-308 (this PR) drops project_id from Task.__init__.
  • src/agent_env/task/task.py:793-797 — put(cls, **kwargs) forwards straight into cls(**kwargs).

Because put passes kwargs through unfiltered, an out-of-repo Task.put(id=…, steps=…, project_id=…) goes from accepted to TypeError: __init__() got an unexpected keyword argument 'project_id', and a Task(…, project_id=…) constructor call does the same. Reads of task.project_id become AttributeError. That's a different failure mode from the stored-document path the description already covers ("a stored task document that still carries the key loads fine") — document compat is handled, the programmatic kwarg isn't.

Worth either a line in the break list or a tolerated-and-ignored project_id kwarg for a release, depending on how much out-of-repo construction you expect. Flagging it only because the verification in the description (git grep project_id) is necessarily scoped to this repo, so callers that construct Task from outside it wouldn't show up in it.


Generated by Claude Code

…B metadata

Replace the hardcoded product/customer/team App tags and the pipeline_step/run_id-only sandbox
tags with one generic _attribution_tags. The shared Modal App is tagged once, so it carries only
the deployment's [sandbox.attribution]; each sandbox carries the run's full attribution filled
from those defaults. E2B metadata is the same resolved dict. A key Modal can't take raises
instead of being rewritten, and the sandbox-started log nests the tags so a key like `name`
can't clobber a LogRecord field.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment on lines +146 to +150
metadata = {
key: value
for key, value in apply_default_attribution(dict(attribution or {})).items()
if value is not None
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Unexposed ports return on reconnect

If a caller sets agent_env_exposed_ports as an attribution key and creates an E2B sandbox with no exposed ports, this code keeps the caller’s value in the sandbox metadata. On reconnect, the provider reads it as its saved port list and adds those ports to tunnel_urls. Code that relies on that list can then treat ports the caller never exposed as available. Keep this provider-owned key out of attribution metadata.

Suggested change
metadata = {
key: value
for key, value in apply_default_attribution(dict(attribution or {})).items()
if value is not None
}
metadata = {
key: value
for key, value in apply_default_attribution(dict(attribution or {})).items()
if value is not None
}
metadata.pop(_EXPOSED_PORTS_METADATA_KEY, None)
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/providers/sandbox_providers/e2b/provider.py
Line: 146-150

Comment:
**Unexposed ports return on reconnect**

If a caller sets `agent_env_exposed_ports` as an attribution key and creates an E2B sandbox with no exposed ports, this code keeps the caller’s value in the sandbox metadata. On reconnect, the provider reads it as its saved port list and adds those ports to `tunnel_urls`. Code that relies on that list can then treat ports the caller never exposed as available. Keep this provider-owned key out of attribution metadata.

```suggestion
        metadata = {
            key: value
            for key, value in apply_default_attribution(dict(attribution or {})).items()
            if value is not None
        }
        metadata.pop(_EXPOSED_PORTS_METADATA_KEY, None)
```

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants