feat(a2a): agents the object store's grants can't reach get their objects through their own staging - #46
Conversation
…ve objects through When an object store's grants cannot reach an agent, as with a local store and an agent on a remote sandbox, agent-env needs another way to hand the agent its objects and collect what it writes. Every SDK agent now serves urn:agentenv:staging/v1 at /ext/staging: a small object store on the agent's own server that agent-env pushes into before a call and pulls from after it. The grants it sends stay ordinary HTTPS grants naming staged paths on the agent's own URL, so extension handlers and the transfer helpers are unchanged. - PUT/GET/DELETE by path, multipart POST under a prefix (the upload-policy shape), and GET/DELETE of a prefix to list or clear it. - Writes land in a staging file and are renamed into place; each write gets a fresh ETag, and DELETE with If-Match spares an object rewritten since it was read. - A path's first segment must be a long, unguessable id; everything staged counts against AGENTENV_STAGING_MAX_BYTES (16 GiB; 0 turns staging off). - python-multipart joins the agent extra for the upload forms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lyan/staged-transfers
…lyan/staged-transfers
…move objects through grant URLs themselves Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…jects through their own staging With a local object store, an agent on a remote sandbox could not reach the store's grants, so skill bundles, snapshots and changelogs were refused and trajectories came back inline. An agent that serves the staging extension (every SDK agent does) now gets ordinary HTTPS grants naming paths on its own staging routes instead, and agent-env moves the bytes over connections it opens itself: - before a call, each object the agent will read is pushed into its staging; - after the call, each object it wrote is pulled into the store, and the call's staging is cleared either way; - a changelog namespace stays staged for the agent's life: its increments are drained into the store every few seconds while a prompt runs, when it ends, and before teardown_run or a teardown_sandboxes step takes the agent down. Every call site that builds a transfer call takes its grants from transfer_store(), which keeps the configured store whenever its grants reach the agent, so hosted stores and local agents are unchanged. Staging errors never name a staged path, and its requests are logged by origin alone. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…its own, and leaves agents its paths when off - The directories holding staged objects are made 0700, so another user on the host can't list the ids that keep them private. A directory given with AGENTENV_STAGING_DIR is used as it is; only the ones staging makes change. - Without AGENTENV_STAGING_DIR each server stages in a directory of its own, made on first use, so a restarted server never counts what an earlier one left against its limit. One server process owns a staging directory: the byte count and the If-Match check are its own. - With staging off, an agent's own route below /ext/staging is no conflict. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lyan/staged-transfers
…ent replaces its stored copy, and last drains are bounded - Each staged write keeps the max_bytes its grant allowed, and a pull stops past it, so an agent whose own client ignores the limit can't fill the host's disk; a changelog namespace keeps its per-object limit the same way. - An increment the agent rewrites after its first copy was drained replaces the stored copy, unless the two hold the same bytes, as the latest write wins through a grant. - The drain at the end of a prompt and the one in a teardown_sandboxes step give up after LAST_DRAIN_SECONDS, so a stalled agent can't hold either up. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… reported once, not every few seconds Found killing an agent's sandbox mid-prompt: the drain loop warned every 5 seconds until the prompt timed out. It now warns when draining starts to fail and logs when it recovers; the last drain still reports what it left. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tten increments are replaced one at a time - While a prompt runs, a failing drain is reported when its reason changes, not just when it starts, and a transport failure reads "the agent could not be reached" whatever httpx's exception, so the reason doesn't flap. - The store overwrites only from memory, so replacing a rewritten increment goes one at a time: at most one is held at once. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lyan/staged-transfers
End-to-end and chaos resultsSetup.
The first pass ran with the agent's documents in a local store while the stage's document database was unreachable. This pass re-ran everything against the stage itself.
Fixed while running these:
Follow-ups, outside this change:
🤖 Generated with Claude Code |
…re refused before any is sent Found in the dev chaos run: with a 1 MB limit and a 4 MB skill, the push was sent anyway and took two minutes to fail, as the agent refused the body before reading it and the client retried. The staging card entry already advertises max_bytes, so a push whose objects together exceed it now fails at once with the same "its staging is full" error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| async def push(self) -> None: | ||
| """Put each object the call reads into the agent's staging; objects that together outgrow the limit | ||
| its card advertises are refused before any is sent.""" | ||
| sizes = await asyncio.gather(*(asyncio.to_thread(_size, self.store, staged.object_url) for staged in self._reads)) |
There was a problem hiding this comment.
Updated objects break staged uploads
If another task replaces a skill or snapshot object while push() is running, _push() can open the new bytes but send the size read earlier as Content-Length. The upload then fails even though the object is available. The new preflight widens this gap by reading every size before any upload starts. Use a size tied to the bytes being sent.
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/a2a_agent/staging.py
Line: 148
Comment:
**Updated objects break staged uploads**
If another task replaces a skill or snapshot object while `push()` is running, `_push()` can open the new bytes but send the size read earlier as `Content-Length`. The upload then fails even though the object is available. The new preflight widens this gap by reading every size before any upload starts. Use a size tied to the bytes being sent.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| async def push(self) -> None: | ||
| """Put each object the call reads into the agent's staging; objects that together outgrow the limit | ||
| its card advertises are refused before any is sent.""" | ||
| sizes = await asyncio.gather(*(asyncio.to_thread(_size, self.store, staged.object_url) for staged in self._reads)) |
There was a problem hiding this comment.
Skill transfers repeat store lookups
push() now starts another metadata lookup for every object, all at once. Grant creation already reads that metadata, and a valid skill bundle can contain 1,000 files. This adds a burst of store requests before any file moves, increasing cost and risking slower transfers or request limits. Reuse the sizes already read or limit the new lookups.
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/a2a_agent/staging.py
Line: 148
Comment:
**Skill transfers repeat store lookups**
`push()` now starts another metadata lookup for every object, all at once. Grant creation already reads that metadata, and a valid skill bundle can contain 1,000 files. This adds a burst of store requests before any file moves, increasing cost and risking slower transfers or request limits. Reuse the sizes already read or limit the new lookups.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
…ts closes it deploy_agent records an agent in the context only once it is configured, and a run's teardown finds only what the context records. So an agent whose snapshot load, changelog apply, changelog enable, MCP registration or skill registration failed kept its sandbox until the provider's TTL. Found in the dev chaos run, where two 100 MB snapshot loads over a slow uplink failed and left their Modal sandboxes running. The step now closes the agent it deployed when anything after the deploy raises, cancellation included; an agent placed on a linked sandbox leaves that sandbox alone, as A2AAgent.close already does. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re left open Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lyan/staged-transfers
What and why
With the local object store, an agent on a remote sandbox (Modal, E2B, a plugin provider) can't reach the store's grants. Today agent-env therefore refuses that agent's skill bundles, snapshots and changelogs ("grants do not reach agents on the 'modal' sandbox provider"), and its trajectories come back inline. This PR makes those flows work with no setup beyond the provider's own credentials: no bucket, tunnel or inbound connection to the user's machine.
An agent that serves the staging extension (#39; every SDK agent does) gets ordinary HTTPS grants naming paths on its own
/ext/stagingroutes. agent-env moves the bytes over connections it opens itself:teardown_runor ateardown_sandboxesstep takes the agent down.If-Match).How it fits in:
transfer_store(). That function keeps the configured store whenever the store's grants reach the agent, so hosted stores, and the local store with the local provider, are unchanged. Staging is chosen only when the grants don't reach, the agent advertises staging, and the agent's URL is HTTPS.StagingErrors that never name a staged path, because the path's id is what keeps it private.Stacked on #32 (the grant-lifetime PR), which it builds on through #13's
grants_reachandsandbox_typeplumbing. It also includes #39's commits, because it uses the staging extension. Merge after #39 and #32, then retarget this PR tomain.How it was tested
pytest tst/unit packages/agentenv-protocol/tests: 5927 passed.tst/unit/a2a_agent/staging_test.pyis new. It uses an in-process SDK agent at an HTTPS URL, so the agent's own calls to its staging run too, and it covers:teardown_rundraining first;-m 'not int_test_slow'): 143 passed. The slow tier runs in CI.🤖 Generated with Claude Code
The PR is not ready to merge while staged uploads can fail when an object changes during transfer.
Fix with agent prompt
Summary
Agents that cannot reach the configured object store can now move object data through staging on their own HTTPS server. Agent-env also drains staged changelogs while prompts run and before agents are torn down.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR Store[Object store] -->|Push reads| Staging[Agent staging] Staging -->|Grants on agent URL| Agent[Agent call] Agent -->|Write results and changelogs| Staging Staging -->|Pull writes and drain increments| StoreReviews (6) · Last reviewed commit: "Merge branch 'edgararakelyan/agent-stagi..."