Skip to content

Latest commit

 

History

History
289 lines (241 loc) · 14.7 KB

File metadata and controls

289 lines (241 loc) · 14.7 KB

Runtime External Payload Transport

External payload storage keeps large encoded workflow payloads outside worker and history JSON. The namespace runtime owns the backing storage provider, credentials, integrity checks, retention, and deletion. Applications use their normal runtime URL, namespace, and role credential; they do not configure or parse the runtime's bucket, container, filesystem path, or provider URI.

Discovery and reference shape

GET /api/cluster/info publishes the active namespace policy at namespace.external_payload_storage and the same transport capability at worker_protocol.server_capabilities.runtime_external_payload_transport. Discovery includes the inline threshold, transport paths, maximum external payload size, timeout budget, expiry policy, cache behavior, typed outcomes, and direct-adapter policy. It intentionally omits backing-driver identity and configuration.

The only remote reference schema is durable-workflow.v2.runtime-external-payload-reference.v1:

{
  "schema": "durable-workflow.v2.runtime-external-payload-reference.v1",
  "reference_id": "ep_01J...",
  "codec": "avro",
  "size_bytes": 4194304,
  "sha256": "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"
}

Payload envelope fields use {codec, external_payload}. Persisted fields named payload_reference use the reference object directly. Provider-specific external_storage envelopes and string provider URIs are not accepted on the remote boundary.

Upload and fetch

Upload encoded bytes with POST /api/external-payloads/v1 using application/octet-stream and these declared metadata headers:

  • X-Durable-Workflow-Payload-Codec
  • X-Durable-Workflow-Payload-Size
  • X-Durable-Workflow-Payload-SHA256

The runtime rejects the request before committing a reference when the declared size exceeds its configured maximum or the observed size or SHA-256 differs. Successful content-addressed retries in the same namespace return the same stable reference identity.

Fetch bytes with GET /api/external-payloads/v1/{referenceId} and send the same metadata headers from the reference. The runtime binds the lookup to the authenticated namespace, fetches the backing bytes, verifies size and SHA-256, and returns application/octet-stream with the verified metadata headers. SDKs verify the returned bytes again before Avro decode.

Both HTTP operations use chunked copying into a verified temporary snapshot, spilling to the process temporary directory above 2 MiB. Provide temporary storage for concurrent transfers as well as the persistent backing store; memory-backed temporary directories still count toward container memory limits. The runtime checks observed bytes even without Content-Length. Ordinary API bodies keep their separate, smaller request limit. Discovery publishes the request timeout. Fetch responses use private, short-lived, immutable caching; SDK caches must be bounded and cannot delete runtime-owned objects.

Workflow descriptions retain the complete input_envelope and output_envelope. For external objects larger than the ordinary request limit, the convenient decoded input or output preview is null; payload_previews.input_omitted and output_omitted distinguish this from a real null value, and max_encoded_bytes gives the preview ceiling. Command diagnostics similarly report payload_preview_omitted when payload-derived fields are omitted. Use the authenticated payload endpoint or SDK to read the complete value. A status lookup must not decode a large object just to display its workflow metadata.

Workflow and activity results are validated before their stored reference is retained. Validation uses Apache Avro schema traversal and the strict Value decoder, without constructing a decoded collection or loading the complete base64 payload. It needs a temporary decoded stream in addition to the verified encoded snapshot. A text scalar is checked for UTF-8 validity and may occupy memory up to its decoded length; byte scalars are skipped after length checks. Size temporary storage and request concurrency together. Metadata projections and standalone activity inspection do not fetch payload bytes.

Self-hosted backing storage

Node-local storage is suitable for a single-node development runtime. A multi-node runtime must place external payload bytes on storage reachable from every HTTP, queue, scheduler, and maintenance process. The published Server image includes an S3-compatible adapter for that purpose.

Set the following environment on every Server process:

DW_EXTERNAL_PAYLOAD_S3_ACCESS_KEY_ID=example-access-key
DW_EXTERNAL_PAYLOAD_S3_SECRET_ACCESS_KEY=example-secret-key
DW_EXTERNAL_PAYLOAD_S3_REGION=us-east-1
DW_EXTERNAL_PAYLOAD_S3_BUCKET=durable-workflow-payloads
DW_EXTERNAL_PAYLOAD_S3_ENDPOINT=https://objects.example.com
DW_EXTERNAL_PAYLOAD_S3_USE_PATH_STYLE_ENDPOINT=false

DW_EXTERNAL_PAYLOAD_S3_ENDPOINT is optional for AWS S3. Access key and secret may both be omitted when the Server workload receives credentials from an IAM role. Temporary credentials can also set DW_EXTERNAL_PAYLOAD_S3_SESSION_TOKEN. Custom S3-compatible services commonly require their HTTPS endpoint and path-style addressing.

Enable the policy with an administrator credential after the bucket exists:

curl -X PUT http://localhost:8080/api/namespaces/default/external-storage \
  -H "Authorization: Bearer $DW_ADMIN_TOKEN" \
  -H "X-Durable-Workflow-Control-Plane-Version: 2" \
  -H "Content-Type: application/json" \
  --data '{
    "driver": "s3",
    "enabled": true,
    "threshold_bytes": 2097152,
    "config": {"prefix": "namespaces/default/"}
  }'

The namespace policy contains no object-store credential. Server resolves the fixed external-payload-s3 disk from process configuration and uses the configured bucket when the policy omits one. If a policy names a bucket, it must match the configured disk bucket.

Before serving workload traffic, call POST /api/storage/test for the namespace. GET /api/cluster/info reports driver_unavailable and a bounded configuration_error such as s3_bucket_missing, s3_credentials_incomplete, or s3_bucket_mismatch without returning credentials, endpoints, bucket names, or object references. Production operators remain responsible for object-store durability, access policy, backup, retention, and recovery validation.

State, expiry, and retention

An upload starts as unclaimed and expires after the advertised abandoned-upload window. The first verified state-bearing request claims it. Claimed references do not expire independently of retained workflow state. Workflow history, schedules, updates, stream items, activity state, and results remain the retention authority. SDKs have no delete endpoint, and namespace/run cleanup does not remove a backing object while retained state in any namespace still owns it.

The singleton maintenance runner executes external-payloads:cleanup --limit=100 on every maintenance pass. The Docker topology runs a pass every 10 seconds and the default Kubernetes CronJob runs a pass every minute when its selected Server image provides the reclamation command. The Kubernetes runner capability-checks the command so a separately qualified older onboarding image remains compatible. Each pass reclaims at most 100 expired references (operators may request up to the hard 1,000-reference ceiling), so cleanup work is bounded and sustained backlog drains over repeated passes. With no blocked storage operation, an expired reference at position p in the oldest-first backlog is attempted within ceil(p / batch-size) maintenance intervals. Cleanup locks the stable backing-object identity, then re-checks expiry and retained ownership before it deletes bytes. References in another namespace keep shared content-addressed bytes alive.

Storage locations are registered in a recoverable writing state before the provider write begins. A successful write promotes the row to ready; a crash or failed final registration leaves the location discoverable and eligible for the same expiry cleanup instead of creating an untracked object.

Operators can inspect namespace-scoped backlog, cumulative deleted counts, blocked outcomes, and storage-driver failures through GET /api/system/external-payload-cleanup or the runtime_external_payload_cleanup section of operator metrics. A blocked pass can be retried with POST /api/system/external-payload-cleanup/pass or the Artisan command. These diagnostics contain aggregate counts only; they do not include provider credentials, provider locations, or reusable reference IDs.

The runtime records upload, fetch, claim, and rejection audit events without logging provider locations, object-store credentials, bearer tokens, or the reusable opaque reference identity. Audit correlation uses a one-way reference identity digest.

Coordinating online backups

A database snapshot does not freeze external objects. Before taking an online backup, an operator can pause Server-owned payload reclamation with a bounded database-backed hold. This is an Artisan-only operation, not a tenant SDK or HTTP API. Upgrade all Server and maintenance processes and apply migrations before relying on it; old processes and direct provider deletions cannot be fenced by this mechanism.

Use a new UUID for each backup operation, and persist it with that operation:

php artisan external-payloads:backup-hold acquire --owner="$BACKUP_ID" --ttl=900
php artisan external-payloads:backup-hold renew --owner="$BACKUP_ID" --ttl=900
php artisan external-payloads:backup-hold status --owner="$BACKUP_ID"
php artisan external-payloads:backup-hold release --owner="$BACKUP_ID"

These commands return JSON. Owner-scoped status exits unsuccessfully when the hold is inactive or belongs to another operation. Unscoped status is for inspection only. Repeating acquire for the active owner returns the existing lease without extending it; use renew explicitly. Expired or released holds must not be revived to authorize an older backup candidate. Release is idempotent for its owner and cannot release a replacement owner's hold.

Acquire waits for an in-flight provider deletion to finish. After it succeeds, start the consistent database snapshot, then copy the external objects while renewing the hold. Normal reads, uploads and workflow execution remain available. Reclamation is deferred for the entire Server, so allow storage headroom for the backup interval. Each lease lasts 1-3600 seconds and renewal cannot extend one operation past 24 hours from acquisition. Database time is the authority, and a lost backup process cannot leave a permanent hold.

Only accept a recovery point after the database and all referenced payload bytes have been durably copied and verified, with the hold still active through that copy. A failed or expired hold invalidates the candidate. Always attempt release in finalization. Use mature database and backup tooling for snapshots, encryption, retention and restore; the hold itself provides none of those. Provider-side changes outside Server, ambiguous provider failures, and an incomplete payload copy still require integrity validation and a failed backup must not replace the last verified recovery point. A restored database may contain the historical hold; release it after validating the restored copy or allow its original deadline to expire.

Payload completion while storage drains

An existing worker claim can retry a draining upload refusal once using the completion schema and header advertised in namespace discovery. Every slot of the same claim shares one byte allowance and a maximum of 128 slots. New client work remains subject to ordinary storage admission.

The prepared local activity candidate on explicit protocol 1.20 additionally advertises completion_context.prepared_schema as durable-workflow.v2.payload-completion-context.v2. Its exact header object contains schema, kind: workflow, task_id, the original workflow attempt, lease_owner, operation and slot, plus one operation identity:

Operation Identity Payload slot
local_activity_checkpoint checkpoint_id Ordinary command payload slot
local_activity_prepare sequence ["descriptor", "arguments"]
local_activity_recover sequence ["descriptor", "arguments"]
local_activity_outcome activity_attempt_id ["report", "result"]

All prepared operations and legacy completion uploads share the original workflow claim's allowance. A new sequence, checkpoint or activity attempt does not reset it. Prepared uploads require the original issued registration, current workflow epoch and live authority. Results additionally require the matching admitted activity attempt and its live execution deadlines. Accepted cancellation fences results from ordinary callbacks. Shielded cleanup keeps the original root deadline.

An upload does not renew workflow or activity leases, application heartbeats, or cancellation deadlines. An exact stored reference can be read after the claim closes without rewriting its object or extending retention. Changed bytes still require live authority. Fenced or stale storage admission refuses both ordinary and completion uploads. Default published protocol 1.19 keeps its existing completion schema.

Typed outcomes and retryability

The transport returns a stable durable-workflow.v2.runtime-external-payload-error.v1 envelope. Outcomes are:

  • external_payload_not_found (404, non-retryable)
  • external_payload_expired (410, non-retryable)
  • external_payload_unauthorized (401 or 403, non-retryable)
  • external_payload_unavailable (503, retryable)
  • external_payload_oversized (413, non-retryable)
  • external_payload_unsupported (415 or 422, non-retryable)
  • external_payload_integrity_mismatch (422, non-retryable)
  • external_payload_namespace_bytes_exhausted (429, retryable)
  • external_payload_namespace_objects_exhausted (429, retryable)
  • external_payload_namespace_quota_unavailable (503, retryable)

Namespace quota rejections include Retry-After and retry_after_seconds. Cumulative usage includes both in-progress and ready objects. Duplicate registration of the same stable object identity does not consume the byte or object budget again.

Reference validation and fetch happen before a workflow/activity completion or control-plane mutation reaches its state transaction. A rejected reference therefore records no partial command and cannot duplicate the intended side effect. Authentication failures use the external-payload unauthorized outcome; a valid credential querying another namespace receives the non-disclosing not-found outcome.

Direct provider adapters may be implemented as an explicitly negotiated self-hosted optimization. They are disabled by default, never required by the standalone Server default or managed Cloud, and do not change the runtime wire reference.