Problem
Every upload always goes through the full joex pipeline (process-item): archive extract → PDF conversion → text extraction/OCR → preview → text analysis. There is no way for the insertion/upload API to say “just store this; processing is not required (or not required immediately).”
That hurts for large reference files — for example Excel workbooks or big merged PDFs — where we mainly need the binary stored and findable as an item, but do not need text extraction or analysis.
Concrete example
Joex spent nearly 3 hours (02:55:11) processing an Excel upload that we only needed stored for reference:

Joex spent nearly 3 hours processing an Excel upload that we only needed stored.
Current behavior
- Upload stores file bytes.
- A
process-item (or multi-upload-process) job is always enqueued.
- The item only becomes UI-visible (
Created) after processing finishes (or fails on last retry).
Closest existing knobs:
- upload
priority (high/low)
skipDuplicates
- global joex OCR / NLP config
None of these provide per-upload store-only.
Desired behavior
A per-upload flag that:
- Creates a visible item + attachments immediately
- Applies given metadata (folder, tags, direction, language, …)
- Skips conversion / OCR / text extraction / preview / analysis
- Allows optional full processing later via the existing reprocess APIs:
POST /api/v1/sec/item/{itemId}/reprocess
POST /api/v1/sec/items/reprocess
Proposed API strawman
Add optional process (boolean, default true) to ItemUploadMeta.
process: true (default) — current behavior
process: false — store-only path: create visible item, skip heavy stages
Available on secured, open/source, and integration upload endpoints; documented in the upload API docs.
Implementation sketch (for discussion)
Keep enqueueing a lightweight process-item job so duplicate-check, CreateItem, and SetGivenData still run, then short-circuit the heavy stages in joex when process: false, and mark the item Created. Reuse existing reprocess for “process later.”
Out of scope (v1)
Related
Happy to implement this if the direction looks good.
Problem
Every upload always goes through the full joex pipeline (
process-item): archive extract → PDF conversion → text extraction/OCR → preview → text analysis. There is no way for the insertion/upload API to say “just store this; processing is not required (or not required immediately).”That hurts for large reference files — for example Excel workbooks or big merged PDFs — where we mainly need the binary stored and findable as an item, but do not need text extraction or analysis.
Concrete example
Joex spent nearly 3 hours (
02:55:11) processing an Excel upload that we only needed stored for reference:Joex spent nearly 3 hours processing an Excel upload that we only needed stored.
Current behavior
process-item(ormulti-upload-process) job is always enqueued.Created) after processing finishes (or fails on last retry).Closest existing knobs:
priority(high/low)skipDuplicatesNone of these provide per-upload store-only.
Desired behavior
A per-upload flag that:
POST /api/v1/sec/item/{itemId}/reprocessPOST /api/v1/sec/items/reprocessProposed API strawman
Add optional
process(boolean, defaulttrue) toItemUploadMeta.process: true(default) — current behaviorprocess: false— store-only path: create visible item, skip heavy stagesAvailable on secured, open/source, and integration upload endpoints; documented in the upload API docs.
Implementation sketch (for discussion)
Keep enqueueing a lightweight
process-itemjob so duplicate-check,CreateItem, andSetGivenDatastill run, then short-circuit the heavy stages in joex whenprocess: false, and mark the itemCreated. Reuse existing reprocess for “process later.”Out of scope (v1)
Related
priority, immediatefileKeyson submitHappy to implement this if the direction looks good.