Skip to content

feat(a2a): agents the object store's grants can't reach get their objects through their own staging - #46

Open
earakely-scale wants to merge 16 commits into
edgararakelyan/store-grant-lifetimefrom
edgararakelyan/staged-transfers
Open

earakely-scale wants to merge 16 commits into
edgararakelyan/store-grant-lifetimefrom
edgararakelyan/staged-transfers

Conversation

@earakely-scale

@earakely-scale earakely-scale commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

What and why

With the local object store, an agent on a remote sandbox (Modal, E2B, a plugin provider) can't reach the store's grants. Today agent-env therefore refuses that agent's skill bundles, snapshots and changelogs ("grants do not reach agents on the 'modal' sandbox provider"), and its trajectories come back inline. This PR makes those flows work with no setup beyond the provider's own credentials: no bucket, tunnel or inbound connection to the user's machine.

An agent that serves the staging extension (#39; every SDK agent does) gets ordinary HTTPS grants naming paths on its own /ext/staging routes. agent-env moves the bytes over connections it opens itself:

  • Reads (skill bundle files, snapshot load, changelog apply): before the call, each object is pushed into the agent's staging, streamed from the store.
  • Writes (trajectory, snapshot save): after the call, each object the agent wrote is pulled into the store. The call's staging is cleared whether the call succeeded or not.
  • Changelog capture: the namespace stays staged for the agent's life. Its increments are drained into the store every 5 seconds while a prompt runs, again when the prompt ends, and before teardown_run or a teardown_sandboxes step takes the agent down.
    • A drained increment is removed from staging only if the agent hasn't rewritten it since (If-Match).
    • If the sandbox dies mid-prompt, at most one drain interval is lost.

How it fits in:

  • Choosing the store. Every call site that builds a transfer call takes its grants from transfer_store(). That function keeps the configured store whenever the store's grants reach the agent, so hosted stores, and the local store with the local provider, are unchanged. Staging is chosen only when the grants don't reach, the agent advertises staging, and the agent's URL is HTTPS.
  • Agents without staging. An agent that doesn't advertise staging gets the inline forms, or the refusal it gets today.
  • Error hygiene.
    • Staging failures are StagingErrors that never name a staged path, because the path's id is what keeps it private.
    • Staging requests are logged by origin only.
    • If an agent can't reach its own URL, the error says so.
  • Unchanged. No request or response model changes, and no new config.

Stacked on #32 (the grant-lifetime PR), which it builds on through #13's grants_reach and sandbox_type plumbing. It also includes #39's commits, because it uses the staging extension. Merge after #39 and #32, then retarget this PR to main.

How it was tested

  • pytest tst/unit packages/agentenv-protocol/tests: 5927 passed. tst/unit/a2a_agent/staging_test.py is new. It uses an in-process SDK agent at an HTTPS URL, so the agent's own calls to its staging run too, and it covers:
    • a skill bundle pushed (3 MB file);
    • a snapshot save pulled, then loaded into the agent;
    • a trajectory pulled;
    • a changelog drained, then applied to the agent;
    • mid-prompt drains plus the final drain;
    • teardown_run draining first;
    • store selection (local agent, no staging, non-HTTPS URL);
    • a full staging failing cleanly without naming a staged path.
  • Integration tier on the local defaults (-m 'not int_test_slow'): 143 passed. The slow tier runs in CI.
  • End to end on Modal with the local object store (claude-code-cli with the staging routes): every flow and six chaos scenarios pass; results are in this comment.

🤖 Generated with Claude Code

RetriggerConfidence Score: 4/5

The PR is not ready to merge while staged uploads can fail when an object changes during transfer.

Fix All in CursorFindings

  1. P1 Updated objects break staged uploads ▶
  2. P2 Skill transfers repeat store lookups ▶
Fix with agent prompt
### Issue 1
src/agent_env/a2a_agent/staging.py:undefined-148
If another task replaces a skill or snapshot object while `push()` is running, `_push()` can open the new bytes but send the size read earlier as `Content-Length`. The upload then fails even though the object is available. The new preflight widens this gap by reading every size before any upload starts. Use a size tied to the bytes being sent.

### Issue 2
src/agent_env/a2a_agent/staging.py:undefined-148
`push()` now starts another metadata lookup for every object, all at once. Grant creation already reads that metadata, and a valid skill bundle can contain 1,000 files. This adds a burst of store requests before any file moves, increasing cost and risking slower transfers or request limits. Reuse the sizes already read or limit the new lookups.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

Agents that cannot reach the configured object store can now move object data through staging on their own HTTPS server. Agent-env also drains staged changelogs while prompts run and before agents are torn down.

  • SDK agents serve staging routes unless staging is turned off.
  • Agent-env pushes objects the agent needs before a call and collects objects it writes afterward.
  • Staged changelog increments move to the object store during prompts and before teardown.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Store[Object store] -->|Push reads| Staging[Agent staging]
  Staging -->|Grants on agent URL| Agent[Agent call]
  Agent -->|Write results and changelogs| Staging
  Staging -->|Pull writes and drain increments| Store
Loading

Reviews (6) · Last reviewed commit: "Merge branch 'edgararakelyan/agent-stagi..."

earakely-scale and others added 7 commits October 1, 2026 10:10
…ve objects through

When an object store's grants cannot reach an agent, as with a local store and
an agent on a remote sandbox, agent-env needs another way to hand the agent its
objects and collect what it writes. Every SDK agent now serves
urn:agentenv:staging/v1 at /ext/staging: a small object store on the agent's own
server that agent-env pushes into before a call and pulls from after it. The
grants it sends stay ordinary HTTPS grants naming staged paths on the agent's
own URL, so extension handlers and the transfer helpers are unchanged.

- PUT/GET/DELETE by path, multipart POST under a prefix (the upload-policy
  shape), and GET/DELETE of a prefix to list or clear it.
- Writes land in a staging file and are renamed into place; each write gets a
  fresh ETag, and DELETE with If-Match spares an object rewritten since it was
  read.
- A path's first segment must be a long, unguessable id; everything staged
  counts against AGENTENV_STAGING_MAX_BYTES (16 GiB; 0 turns staging off).
- python-multipart joins the agent extra for the upload forms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…move objects through grant URLs themselves

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…jects through their own staging

With a local object store, an agent on a remote sandbox could not reach the
store's grants, so skill bundles, snapshots and changelogs were refused and
trajectories came back inline. An agent that serves the staging extension
(every SDK agent does) now gets ordinary HTTPS grants naming paths on its own
staging routes instead, and agent-env moves the bytes over connections it
opens itself:

- before a call, each object the agent will read is pushed into its staging;
- after the call, each object it wrote is pulled into the store, and the
  call's staging is cleared either way;
- a changelog namespace stays staged for the agent's life: its increments are
  drained into the store every few seconds while a prompt runs, when it ends,
  and before teardown_run or a teardown_sandboxes step takes the agent down.

Every call site that builds a transfer call takes its grants from
transfer_store(), which keeps the configured store whenever its grants reach
the agent, so hosted stores and local agents are unchanged. Staging errors
never name a staged path, and its requests are logged by origin alone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…its own, and leaves agents its paths when off

- The directories holding staged objects are made 0700, so another user on the
  host can't list the ids that keep them private. A directory given with
  AGENTENV_STAGING_DIR is used as it is; only the ones staging makes change.
- Without AGENTENV_STAGING_DIR each server stages in a directory of its own,
  made on first use, so a restarted server never counts what an earlier one
  left against its limit. One server process owns a staging directory: the
  byte count and the If-Match check are its own.
- With staging off, an agent's own route below /ext/staging is no conflict.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale
earakely-scale requested a review from a team as a code owner October 1, 2026 23:18
Comment thread src/agent_env/a2a_agent/staging.py
Comment thread src/agent_env/a2a_agent/object_transfer.py
Comment thread src/agent_env/a2a_agent/staging.py
Comment thread src/agent_env/task_step/task_steps/teardown_sandboxes.py
…ent replaces its stored copy, and last drains are bounded

- Each staged write keeps the max_bytes its grant allowed, and a pull stops
  past it, so an agent whose own client ignores the limit can't fill the
  host's disk; a changelog namespace keeps its per-object limit the same way.
- An increment the agent rewrites after its first copy was drained replaces
  the stored copy, unless the two hold the same bytes, as the latest write
  wins through a grant.
- The drain at the end of a prompt and the one in a teardown_sandboxes step
  give up after LAST_DRAIN_SECONDS, so a stalled agent can't hold either up.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/agent_env/a2a_agent/staging.py Outdated
… reported once, not every few seconds

Found killing an agent's sandbox mid-prompt: the drain loop warned every
5 seconds until the prompt timed out. It now warns when draining starts to
fail and logs when it recovers; the last drain still reports what it left.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/agent_env/a2a_agent/staging.py Outdated
earakely-scale and others added 3 commits October 1, 2026 14:39
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tten increments are replaced one at a time

- While a prompt runs, a failing drain is reported when its reason changes,
  not just when it starts, and a transport failure reads "the agent could not
  be reached" whatever httpx's exception, so the reason doesn't flap.
- The store overwrites only from memory, so replacing a rewritten increment
  goes one at a time: at most one is held at once.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@earakely-scale

earakely-scale commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

End-to-end and chaos results

Setup.

  • agent-env from this branch, on a laptop.
  • claude-code-cli built with the staging routes, registered in a development stage (document store, image registry and bucket), running on Modal.
  • The run's objects in the local object store, so every grant goes through the agent's staging.
  • Plain task runs, with the configured object store local, and agent-env run bundles (an @local run whose writes route to the local store).

The first pass ran with the agent's documents in a local store while the stage's document database was unreachable. This pass re-ran everything against the stage itself.

Flow Result
Skill bundle, 4 MB, pushed into staging ✅ The agent answered with the code word only a file in the bundle held.
Changelog capture ✅ 5 increments, drained mid-prompt and at the prompt's end; all valid archives; nothing left staged.
Trajectory, pulled ✅ In the local store (12,908 bytes).
Snapshot save, pulled, then load into a fresh agent ✅ The fresh agent read the code word back.
Rewind a fresh agent through the changelog ✅ In full: both files as written. Up to tool call 4 (exclusive): the last append correctly absent.
The same through an agent-env run bundle ✅ Skill, changelog and trajectory. The snapshot is refused by @local routing, which predates this change: a snapshot artifact's id isn't an @local id.
Chaos scenario Result
Agent sandbox terminated after 4 increments were drained ✅ All 4 are whole archives in the store, with no strays. The drain warned once over the 20 minutes the prompt then waited for its own A2A timeout.
agent-env SIGKILLed with 20 MB of a 300 MB snapshot pull in flight ✅ No workspace object reached the store, and the capture was never registered. A rerun stored all 300,097,706 bytes.
3 changelog runs at once ✅ 3 distinct namespaces; each holds only its own agent's files.
Staging limit (1 MB) below a 4 MB skill ✅ staging an object on the agent failed: its staging is full, in 4 s, with no staged path in any error or metadata.
Snapshot restore beyond the staging limit ✅ The deploy fails in 19 s, and its sandbox is closed.
100 MB workspace ✅ Pulled into the store. ❌ Two attempts to load it back failed over a ~50–120 KB/s uplink: staging an object on the agent failed: the agent could not be reached. On a normal uplink the same load passed in 31 s with the same sha256.
Staged increments deleted before the drain took them ✅ Skipped (a gap, which the contract allows); the run completed, and the other 9 are stored.

Fixed while running these:

  • e15fb36, 7adbb47: a drain that kept failing warned every 5 s; it now reports when its reason changes.
  • a32835d: objects that outgrow the agent's advertised staging limit are refused before any is sent. The push used to take 2 minutes to fail.
  • 0df0dba: a deploy_agent step that fails after its sandbox exists closes it. Before, a failed snapshot load, changelog apply or skill registration left the sandbox running until its TTL, because teardown finds only agents the context records. The two failed 100 MB loads above left two this way; both are now terminated.

Follow-ups, outside this change:

  • Pushes and pulls restart from zero when they fail. A large object over a slow or flaky uplink can't complete. Chunked, resumable transfers would fix that.
  • Throughput is bounded by the provider tunnel's egress. It was 0.2–1.5 MB/s per stream on Modal, against a 4.5 MB/s downlink measured directly. Parallel ranged pulls would help.
  • prompt_agent waits out its full timeout on an agent whose sandbox died.
  • A killed agent-env leaves its pull's temp file in the system temp directory, outside the store.

🤖 Generated with Claude Code

…re refused before any is sent

Found in the dev chaos run: with a 1 MB limit and a 4 MB skill, the push was
sent anyway and took two minutes to fail, as the agent refused the body
before reading it and the client retried. The staging card entry already
advertises max_bytes, so a push whose objects together exceed it now fails
at once with the same "its staging is full" error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
async def push(self) -> None:
"""Put each object the call reads into the agent's staging; objects that together outgrow the limit
its card advertises are refused before any is sent."""
sizes = await asyncio.gather(*(asyncio.to_thread(_size, self.store, staged.object_url) for staged in self._reads))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Updated objects break staged uploads

If another task replaces a skill or snapshot object while push() is running, _push() can open the new bytes but send the size read earlier as Content-Length. The upload then fails even though the object is available. The new preflight widens this gap by reading every size before any upload starts. Use a size tied to the bytes being sent.

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/a2a_agent/staging.py
Line: 148

Comment:
**Updated objects break staged uploads**

If another task replaces a skill or snapshot object while `push()` is running, `_push()` can open the new bytes but send the size read earlier as `Content-Length`. The upload then fails even though the object is available. The new preflight widens this gap by reading every size before any upload starts. Use a size tied to the bytes being sent.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Cursor Fix in Claude Code Fix in Codex

async def push(self) -> None:
"""Put each object the call reads into the agent's staging; objects that together outgrow the limit
its card advertises are refused before any is sent."""
sizes = await asyncio.gather(*(asyncio.to_thread(_size, self.store, staged.object_url) for staged in self._reads))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Skill transfers repeat store lookups

push() now starts another metadata lookup for every object, all at once. Grant creation already reads that metadata, and a valid skill bundle can contain 1,000 files. This adds a burst of store requests before any file moves, increasing cost and risking slower transfers or request limits. Reuse the sizes already read or limit the new lookups.

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/agent_env/a2a_agent/staging.py
Line: 148

Comment:
**Skill transfers repeat store lookups**

`push()` now starts another metadata lookup for every object, all at once. Grant creation already reads that metadata, and a valid skill bundle can contain 1,000 files. This adds a burst of store requests before any file moves, increasing cost and risking slower transfers or request limits. Reuse the sizes already read or limit the new lookups.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Cursor Fix in Claude Code Fix in Codex

earakely-scale and others added 3 commits October 2, 2026 09:29
…ts closes it

deploy_agent records an agent in the context only once it is configured, and
a run's teardown finds only what the context records. So an agent whose
snapshot load, changelog apply, changelog enable, MCP registration or skill
registration failed kept its sandbox until the provider's TTL. Found in the
dev chaos run, where two 100 MB snapshot loads over a slow uplink failed and
left their Modal sandboxes running. The step now closes the agent it deployed
when anything after the deploy raises, cancellation included; an agent placed
on a linked sandbox leaves that sandbox alone, as A2AAgent.close already does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…re left open

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant