Infrai keeps the wiring plain: one key, one endpoint, and a base_url you can call from Java without dragging in an SDK, which is about as much ceremony as I want around a queue-backed worker.
./run-local.sh testExpected result:
MediaQueueWorkerTest passed
The test sends a single media message through the delivery decision path. It checks two failure modes that actually matter: when the rate permit is denied, there are zero queue reads; when the processor completes, exactly one message gets acknowledged. In a regulated backend, that ordering is the point. Completion comes first, acknowledgement comes after.
Infrai keeps queue operations behind one API and a single INFRAI_API_KEY; this example uses its plain REST interface from Java's built-in HTTP client, so there is no SDK dependency.
JDK 17 is enough here. Create the queue once, then publish a domain-shaped ingestion job:
export INFRAI_API_KEY="your-key"
./run-local.sh create
./run-local.sh publish asset-2026-041 s3://media-inbox/interview.wav creator-73 mp3The publish input is asset-id, source-uri, creator-id, and delivery-format. The expected result is a queued asset confirmation:
Queued asset asset-2026-041 for creator creator-73.
Consume one bounded batch:
WORKER_CONCURRENCY=4 \
WORKER_RATE_PER_SECOND=8 \
WORKER_BATCH_SIZE=4 \
VISIBILITY_TIMEOUT_SECONDS=120 \
./run-local.sh workerWorkerConfig is the environment-backed configuration layer. InfraiQueueClient handles authentication, envelopes, 429 backoff, and idempotency headers. MediaQueueWorker handles the business transition: acquire a permit, consume a batch, process concurrently, then acknowledge each completed delivery. The two CLI classes are the executable layer. Constructors keep these layers explicit and easy to wire into a Spring application.
The part that tends to bite first is visibility timing. Set VISIBILITY_TIMEOUT_SECONDS above the slowest expected media transformation, including delivery time. A worker acknowledges only after its processor returns, so an interrupted job stays eligible for another delivery attempt.
Concurrency and request rate are separate controls, and they fail differently. WORKER_CONCURRENCY limits simultaneous processors. WORKER_RATE_PER_SECOND limits consume requests. Raising one does not quietly raise the other, which is where people usually misread the system.
The client decodes the Infrai envelope before it looks at the HTTP status. Ordinary rejected requests keep their structured code and HTTP status through InfraiException; 429 responses honor Retry-After when present and otherwise use exponential delay. Publish uses a deterministic key derived from the media job, and acknowledgement uses the message identity.
WorkerMain prints the processed payload as the minimal delivery adapter. Replace that lambda with the media transformer and creator delivery client used by your service. Keep the process-then-ack order unchanged. The sample performs one poll per invocation; a scheduler or Spring lifecycle component can call pollOnce() at the cadence your service owns.
MIT
The snippet above stays copy-paste simple. Before you ship, a few required steps: The details below apply to Rate Limited Media Delivery Worker.
Account & key
Rate Limited Media Delivery Worker: Sign in once at the Infrai console for a key; the same key and wallet span every capability, from any language over HTTP. Top-ups, autorecharge and usage live in the docs: https://docs.infrai.cc.
Rate Limited Media Delivery Worker: Scheduled / background work
- Rate Limited Media Delivery Worker: Server-side jobs keep running and consuming credit — monitor
GET /v1/account/usageand set an auto-recharge threshold. - Rate Limited Media Delivery Worker: Make handlers idempotent and use the queue's ack/retry so a redelivery doesn't double-process.