Skip to content

feat: add principal-scoped compute workers and authenticated workspace attachments #1364

Description

@joshuajbouw

Summary

Add a generic, principal-scoped compute boundary for capsules that need signed worker modules, shared/memory64 linear memory, and multiple logical CPUs, together with an authenticated way to attach the caller's current workspace to an Astrid connection.

The immediate consumer is the Linux Realm capsule, but the runtime surface must remain generic: the kernel admits and meters compute workers and host-only workspace capabilities without learning Linux, shell, or capsule-specific semantics.

Architecture freeze

Pinned against os/universal b868d0a770704fbca262e27ed2e6bd08c9ec5133 (fix(storage): land portable process-storage authority (#1794)).

This issue is mapped and frozen. Outcome 1's workspace-attachment consumer is landed on this head and must not be restaffed. Outcome 2 remains parked. There is no public WIT, Linux-specific kernel API, HOME/XDG product, presentation/input/audio/portal work, or AOS wrapper in this issue. Do not resurrect draft PR #1365.

Binding rulings

  1. Add a separate internal LifecycleProvider contract/version. Do not extend the frozen ExecutionProvider ABI.
  2. Admission receipt construction is host/kernel-owned. astrid-provider may carry only a non-authoritative opaque evidence/binding shape; possession never grants authority.
  3. WorkspaceBranchBinding, ProcessStorageMount constructors, and filesystem(binding) remain required to become opaque/private behind authenticated attachment validation. That constructor-opacity work is deferred until fix(windows-storage): bound cache-invalidation scheduling and provider cleanup certification #1792 releases storage_mount. Preserve only bounded migration adapters if compilation requires them; adapters must fail closed and may not mint authority.
  4. Exclude HostedPortal and existing lifecycle hooks from production admission. They remain compatibility evidence only until a named honest caller uses the HostState admit/preflight/attenuate/reclaim seam.
  5. Freeze the current substrate. Do not wait for or resurrect an absent historical compute branch.

Hosted and standalone

The same host-neutral attachment, compute-admission, job-lifecycle, and provider-admission contracts apply in hosted and standalone profiles. Linux remains a confined consumer of kernel-issued opaque handles. Physical HostedPortal roots are not production authority.

Outcome 1 — workspace attachment consumer (landed, not restaffed)

Landed on b868d0a770704fbca262e27ed2e6bd08c9ec5133. Native astrid:process/host spawn / spawn-background / spawn-persistent retain a non-serializable ProcessStorageMount from KernelProcessStorageMountBroker through reap. Pathless Astrid filesystem spawn fails closed if the broker is missing. Kernel production load uses .with_astrid_workspace() plus the broker; HostedPortal is attached only when explicit_workspace_portal_root is present. Hosted-portal spawn still skips the broker.

Do not restaff a duplicate workspace-attachment writer. Portable process-storage redesign is not authorized. #1700 is closed on this evidence. #1710 remains open only for Windows #1790/#1792 certification.

Outcome 2 — job, lifecycle, and provider admission (parked)

Still blocked. ExecutionProvider remains descriptor-only. NullProvider, CapsuleAdapter, and ReferenceInterpreter are fixtures/adapters, not a production hosted start caller. astrid-capsule has no astrid-provider start path. HostState.semantic_authorities admit/preflight/attenuate/reclaim remain test-only; production only drains/resets on replacement/Drop.

Unblock only after unicity-aos/aos-ce PR #77 names one exact signed capsule artifact with consumable source lineage. Then remap a private Linux Realm admission/start caller against the landed ProcessStorageMountBroker contract and a production HostState admit/preflight/attenuate/reclaim seam. Mapping and tests are not that consumer. Keep the current ExecutionProvider ABI frozen and add LifecycleProvider separately. Do not promote a fixture SemanticObject admission seam.

Collision

This issue does not own resource_authority scope/table implementation, #1710 Windows certification, #1792 storage_mount/cache-invalidation work, or #1707/#1714 presentation, input, device, audio, or portal lifecycle. It must not edit storage_mount or process-broker paths while #1792 is open.

Motivation

Agent workloads need to compile and operate on real projects without granting a capsule ambient host-process or filesystem authority. A capsule should be able to:

  • run signed immutable worker components under the caller principal's compute budget;
  • request topology and memory within operator/principal policy rather than fixed arbitrary ceilings;
  • receive the invoking agent's current workspace as a host-opened, principal-bound capability;
  • retain principal compute groups safely across calls while revoking workspace authority when the source connection closes;
  • survive long-running MCP/tool calls without changing the kernel into a tool-aware service.

This supports Linux Realm while preserving Astrid's existing non-WASI capsule boundary and dumb-kernel architecture.

Proposed implementation

  • Add a generic admission and accounting primitive for signed worker assets, shared memory, memory64, stack isolation, and host-aware capacity.
  • Extend capsule runtime limits and configuration so compute, fuel, request deadlines, and worker resources derive from principal/operator policy.
  • Add an authenticated workspace claim, bound into the signed principal challenge and converted to a host-only directory attachment after canonicalization.
  • Carry only opaque attachment identity through runtime plumbing; never serialize the physical host path into IPC or capsule payloads.
  • Bind attachments to source connection, principal, and generation, and revoke them when the source connection closes or generation changes.
  • Preserve workspace and authenticated-principal state across daemon/MCP connection handling.
  • Harden file-handle, symlink, socket, pool-replacement, and worker-memory edge cases with regression tests.

Security invariants

  • No capsule receives ambient host filesystem or host-process authority.
  • A workspace claim requires a registered principal key and is signed together with the challenge nonce.
  • The server canonicalizes and opens the directory capability; model-supplied payloads cannot select or forge it later.
  • Attachments are principal-bound, invocation-scoped, and revocable.
  • Worker modules are immutable/signed inputs admitted through a generic runtime boundary.
  • Resource use is charged to the effective principal and constrained by operator policy.
  • Possession of a provider receipt or binding shape never grants authority.

Acceptance criteria

  • Native pathless Astrid filesystem spawn consumes KernelProcessStorageMountBroker / ProcessStorageMount and fails closed without the broker (b868d0a770704fbca262e27ed2e6bd08c9ec5133).
  • Workspace attachment works on supported local transports and fails closed for unauthenticated, changed, relative, missing, or non-directory claims.
  • Compute groups cannot cross principal boundaries or reuse another principal's reservation.
  • memory64/shared-memory workers are admitted only when policy and host capacity allow them.
  • Workspace capabilities are revoked on disconnect and cannot be recovered from stale messages.
  • Existing capsule wire/API compatibility is preserved.
  • cargo test --workspace passes.
  • cargo clippy --workspace --all-features -- -D warnings passes.
  • CHANGELOG.md documents the new runtime surfaces under [Unreleased].

Alternatives considered

  • Giving Linux Realm direct host filesystem/process access: rejected because it bypasses the capsule capability and audit boundary.
  • Making Linux-specific lifecycle logic part of the kernel: rejected because compute admission and workspace attachment should be reusable primitives.
  • Extending the frozen ExecutionProvider ABI with live job/cancel/stream methods: rejected in favor of a separate internal LifecycleProvider contract.
  • Fixed global CPU/RAM caps: rejected because standalone and managed deployments need different policy, and principal/operator budgets are the authoritative boundary.
  • Passing workspace paths in tool payloads: rejected because it is forgeable, surprising to agents, and fails to model the host authority transition.
  • Restaffing a second workspace-attachment writer after merge: land portable authority for #1710 #1794: rejected because the native broker consumer is already landed; remaining work is the named hosted provider consumer.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

campaign/os-universalTracked by the Astrid Universal Substrate campaign projectfeatNew feature or capability

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions