Turn invoices, receipts, and forms into structured data. Read a document, extract the fields you define, flag low-confidence values for a human, and run the actions you configured once the data is confirmed.
It is a multi-tenant app you can run yourself. Documents come in by upload or email. A background pipeline runs OCR when needed, sends the text and the original image to Claude with a per-template schema, and gets back typed fields with per-field confidence. Anything below the template's threshold routes to a review queue where a human confirms or corrects it. On confirm, the action engine runs the webhooks, notifications, or record marks you set up.
- A user uploads a document, or forwards one to the workspace's inbound email address.
- The file lands in object storage and a BullMQ job picks it up.
- If it is a scan or image-based PDF, Tesseract.js pulls the raw text via OCR.
- Claude receives the text and the original image plus the template's extraction schema, and returns structured JSON with a confidence score per field.
- Fields below the template's threshold get flagged. If any field is flagged, the whole document goes to the review queue.
- A reviewer claims the task, corrects the flagged fields, and resolves it.
- Once confirmed (auto or human), the action engine fires the template's configured actions.
- The dashboard tracks pipeline status, per-template accuracy, and what is waiting for review.
Dashboard. Upload PDFs, scans, or images. The dashboard shows the whole pipeline: how many documents are confirmed, waiting on review, still processing, or failed, plus per-template accuracy and a recent activity feed.
Split-view extraction. Every document opens to the original file on one side and the extracted fields on the other. Each field carries a confidence badge, green when the model was sure, amber when it was not. The PDF renders inline with zoom and rotate, so you can check the source without leaving the page.
Human review. Low-confidence extractions route to a review queue automatically. A reviewer claims the task and can edit any field. Flagged fields are highlighted, and nested line-item tables are editable row by row. Only the fields that actually changed are saved as corrections, and the reviewer's values are trusted at full confidence.
Templates. Each document type is a template with a typed field schema (string, number, date, currency, and nested lists), a confidence threshold, and a set of actions. The threshold slider previews how many of your real documents would auto-confirm versus route to review before you save.
Actions. When a document is confirmed, its template's actions run: post to a webhook, notify a teammate, or mark a record. Each run is logged against the document.
Multi-tenant. Every organization has its own workspace, documents, templates, review queue, and members with roles. Documents can arrive by upload or by forwarding them to the org's inbound email address; attachments become documents and run the same pipeline.
| Layer | Choice |
|---|---|
| Frontend | Next.js (App Router) + Tailwind CSS + shadcn/ui, dark theme |
| Backend | Bun + Express, 3-layer architecture (routes, controllers, services) |
| Database | PostgreSQL + Drizzle ORM |
| Auth | Better Auth with the organization plugin (workspaces, roles, invites) |
| Queue | BullMQ + Redis, self-hosted, no external queue vendor |
| Object storage | MinIO (S3-compatible), runs locally in Docker |
| OCR | Tesseract.js, self-hosted, no per-page cost |
| Extraction | Claude (anthropic/claude-sonnet-5 via OpenRouter), JSON-schema structured output per template |
| Parsing | unpdf for PDF text, mammoth for DOCX, exceljs for XLSX |
| Testing | Vitest + Testing Library (unit/component), Playwright (e2e), real Postgres + MinIO + Claude integration tests |
unpdfoverpdf-parse. pdf-parse is unmaintained and does synchronous filesystem reads for test fixtures, which breaks in serverless and no-filesystem environments. unpdf is the maintained successor on the same pdf.js core, with better types.exceljsoverxlsx(SheetJS). The free SheetJS npm package has unpatched high-severity vulnerabilities because the maintained version moved behind a paid product. exceljs is the standard open alternative.- BullMQ over Inngest. BullMQ needs no external account or API keys for local development or testing, just a Redis container, and its job and queue model matches this pipeline directly.
- Send the image, not just the OCR text. Tesseract outputs flat text with no table structure. For invoices with line items, docflow sends both the raw OCR text and the original document image to Claude, which reads images natively and reconstructs table structure that flat text alone would lose.
- Structured output per template. Extraction uses a JSON schema derived from each template, so responses match the expected shape instead of relying on prompt instructions alone.
Prerequisites: Bun, Docker (for Postgres, Redis, MinIO), and an OpenRouter API key.
docker compose up -d # postgres, redis, minio
cd backend && bun install
bun run db:migrate # apply the schema
bun run dev # api on :4000
bun run worker # pipeline worker (separate terminal)
cd frontend && bun install
bun run dev # app on :3000Open http://localhost:3000, create an account and an organization, add a template (or use the seeded defaults), and upload an invoice. See .env.example in each service for required environment variables.
docflow/
backend/ # Bun + Express API, BullMQ pipeline, Drizzle schema (src/db)
frontend/ # Next.js app
docker-compose.yml
assets/ # README media
The backend follows a 3-layer structure: api/routes wire middleware to api/controllers, which validate input with Zod and call api/services for the actual work. The document pipeline lives under src/queue as discrete stages (parse, OCR, extract, flag, act).
# backend
cd backend
bun run test # 100+ unit + integration (real Postgres, MinIO, Claude)
bun run test:e2e # playwright
bun run build # tsc --noEmit
bun run lint # biome
# frontend
cd frontend
bun run test # component tests (vitest + testing library)
bun run test:e2e # playwright
bun run build # next build
bun run lint # biomeIntegration tests hit real services, not mocks: real Postgres and MinIO for the pipeline, and real Claude calls for extraction. The inbound-email path is verified end to end through actual HTTP.
Self-hosted. Runs on Docker (Postgres, Redis, MinIO) with an OpenRouter key for extraction. Point a mail provider at the inbound endpoint to enable email ingestion.






