A small, synchronous Python pipeline that fans interview PDFs out through Celery, stores every artifact in RustFS, and uses Redis as the Celery broker/result backend. PDF conversion, OCR, LLM extraction, and client submission are deterministic mocks so the complete flow is runnable before real integrations exist.
- Python
3.14tfree-threaded build (from.python-version) - Celery
5.6.xwith Redis - Redis
7.4(fromcompose.yaml) - RustFS
1.0.0-rc.6, using its S3-compatible API - boto3, Pydantic Settings, and Click
- uv, Ruff, ty, flake8 cognitive complexity, and pytest
make setup
docker compose up --build -d
docker compose run --rm cli seed doc-001
docker compose run --rm cli seed doc-002
docker compose run --rm cli scan
docker compose logs -f workerAfter the worker logs show completion:
docker compose run --rm cli show-result doc-001
docker compose run --rm cli show-result doc-002seed doc-001/seed doc-002 embed built-in interview-record PDFs so they
work inside the Compose image even without host files. For the fuller two-page
versions kept in the repo, run from the host checkout instead:
uv run pipeline seed doc-001 --file demo/pdfs/doc-001.pdf
uv run pipeline seed doc-002 --file demo/pdfs/doc-002.pdfThe demo PDFs live in demo/pdfs/ and are regenerated by
uv run python scripts/make_demo_pdfs.py.
RustFS is exposed at http://localhost:9000; its console is at
http://localhost:9001. make secret generates the local credentials in
.env; no credential is committed.
For each <document_id>/data.pdf, the worker writes:
<document_id>/imgs/1.png
<document_id>/ocr/1.json
<document_id>/structured-text.txt
<document_id>/result
A submitted document also gets <document_id>/submission.json, which records
the mock client receipt.
scan only queues documents that have data.pdf and do not yet have result.
It queues submit_to_process, which validates the ID and fans out the process
task. If PIPELINE_AUTO_SUBMIT=true, processing queues submit_to_client
automatically after writing the result.
Keep PIPELINE_AUTO_SUBMIT=false for human review. Submit selected IDs through
Celery after checking their results:
docker compose run --rm cli submit-to-client doc-001 doc-002Use --now to invoke the same underlying workflow synchronously without the
queue:
docker compose run --rm cli process doc-001
docker compose run --rm cli submit-to-client --now doc-001To watch the objects appear in RustFS while the worker runs, open the RustFS
console at http://localhost:9001 (sign in with the keys in .env) and browse
the interviews bucket. A document-centric debug web UI is planned; see
.scratch/rustfs-debug-console/spec.md.
Every Celery task has a Click command that can exercise its workflow directly.
pipeline scan [--now]
pipeline submit-to-process [--now] DOCUMENT_ID...
pipeline process DOCUMENT_ID
pipeline submit-to-client [--now] DOCUMENT_ID...
pipeline seed [--file PDF] DOCUMENT_ID
pipeline show-result DOCUMENT_ID
Run these locally with uv run pipeline ... or in Compose with
docker compose run --rm cli .... Commands that touch the queue
(scan, submit-to-process, submit-to-client without --now) must use the
Compose variant for the docker stack, because they enqueue to
PIPELINE_REDIS_URL (redis://redis:6379/0 inside Compose) rather than a
host Redis.
- PDF-to-image logic:
src/pipeline_problem/models/document.py - OCR adapter:
src/pipeline_problem/services/ocr.py - LLM adapter:
src/pipeline_problem/services/llm.py - External client adapter:
src/pipeline_problem/services/client.py - RustFS adapter:
src/pipeline_problem/services/rustfs.py - Orchestration:
src/pipeline_problem/workflows.py - Celery task boundaries:
src/pipeline_problem/tasks.py
Keep remote adapters synchronous. Celery provides process-level concurrency, so application code does not need asyncio or custom threads.
make lint
make test
make audit
make format
make cimake down stops the stack. To also remove stored demo objects, run
docker compose down -v.